PrismML's Ternary Bonsai 2 Runs 262K Context on a 12 GB GPU
A patched serving stack runs the 27B Ternary Bonsai 2 at the full 262k context window on a single 12 GB consumer GPU, with agentic tool-calling finally fixed.
- New serving stack runs Ternary Bonsai 2 27B at full 262k context on a 12 GB GPU.
- Hits 100 tok/s decode at 32k context, 87 at 64k, 70 at 112k on an RTX 4070.
- Tiered KV cache keeps hot positions in VRAM, overflow in pinned system RAM, bit-identical output.
- Fixes agent breakage: tool calls parse 9 of 9 vs 1 of 9, HumanEval goes from 0 to 160.
- AppWorld agent benchmark climbs from 64.3% to 70.8% with same weights, patched serve.
- Weights byte-identical to PrismML base; MIT serve, Apache 2.0 weights, Windows bundle included.
Ternary Bonsai 2 Gets a 262K Serving Stack for 12 GB GPUs
A community serving package for PrismML’s 27-billion-parameter Ternary Bonsai 2 reports support for the model’s full 262,144-token context window on a 12 GB RTX 4070. It keeps the key-value cache at q8_0 precision, reaches 100 generated tokens per second at 32K context, and adds server-side fixes for reasoning and tool-calling APIs.
The package combines the existing 1.58-bit ternary model with a Multi-Token Prediction draft head and 33 llama.cpp patches. The project attributes several reported failures to server defaults, template validation, and tool formatting. All performance and benchmark figures below come from the publisher.
One model, a rebuilt runtime
The release page distributes GGUF files for llama.cpp-based inference. The underlying model comes from PrismML’s base weights, which encode most parameters with three possible values and average about 1.58 bits per weight.
According to the release, all 851 tensors from the base PTQ1_0 file remain byte-identical. The only model-file additions are 15 tensors for the draft head, which proposes several likely tokens before the main model verifies them. The serving code, launcher, and optional server layer use the MIT license; PrismML’s weights retain the Apache 2.0 license.
Triple-digit decode at 32K
On an RTX 4070 with 12 GB of VRAM, the patched server reports the following generation rates. Decode speed measures token generation after prompt processing, so these figures do not describe the time required to ingest a long prompt.
| Context length | Patched server | Reference server |
|---|---|---|
| 32K |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.