OpenBMB's MiniCPM5-2B Gets a 324M Draft Model That Accepts 6 Tokens at Once
OpenBMB released a 324M-parameter draft model that accelerates MiniCPM5-2B inference using semi-autoregressive speculative decoding, hitting 5.5x accepted length on average.
- OpenBMB released MiniCPM5-2B-DSpark, a 324M-parameter speculative decoding draft model for MiniCPM5-2B.
- Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.
- Math and code domains hit 6.05 and 6.11 accepted tokens per verification step respectively.
- Uses DSpark semi-autoregressive drafting: parallel backbone plus lightweight sequential head to prevent suffix decay.
- Served via SGLang with
--speculative-algorithm DSPARKand block size 7, Apache 2.0 licensed. - Trained on 7.05B tokens, 6 epochs, with cross-entropy, L1, and confidence loss objectives.
OpenBMB adds a 324M DSpark draft to MiniCPM5-2B
OpenBMB has released the DSpark checkpoint, a 323.8M-parameter draft model trained specifically to accelerate MiniCPM5-2B through speculative decoding. SGLang uses the smaller model to propose seven-token blocks, then asks the 2.52B-parameter target to verify those proposals in one pass.
MiniCPM5-2B is a dense next-token model with 42 layers, grouped-query attention, and a native context window of 131,072 tokens. Its LlamaForCausalLM architecture allows established inference engines to load the model without custom attention kernels.
A published benchmark overview reports a 53.9 average across 34 evaluations, compared with 51.1 for Qwen3.5-4B. Reported results include 69.1 on LiveCodeBench v6 and 97.1 on τ2-Bench Telecom. Aggregate scores combine distinct tasks, so deployment tests remain necessary for any specific workload.
Seven-token blocks, one verifier
Speculative decoding lets a smaller model propose upcoming tokens at lower computational cost. The target evaluates the proposed block in parallel, accepts tokens that satisfy its verification rule, and regenerates from the first rejected position. The target remains responsible for the final token sequence, while draft accuracy determines how much work each verification pass completes.
DSpark addresses the acceptance decay common to parallel drafters, whose later proposals often lack enough information about earlier tokens in the same block. Its parallel backbone generates the initial proposals, a lightweight sequential module adds dependencies within the block, and a confidence scheduler chooses how many proposals to verify based on estimated prefix survival and the serving engine’s throughput profile.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.