vLLM Doubles DeepSeek V4.1-Flash Speed With 5x Gains Under Heavy Agent Workloads
vLLM shipped three weeks of optimizations for DeepSeek-V4.1-Flash, cutting latency nearly in half and lifting agentic throughput over 5x.
- vLLM made DeepSeek-V4.1-Flash 1.9x faster at low concurrency and 5.3x higher throughput on AgentX.
- SWA bounded replay reruns only the last 128 tokens, skipping nearly half the model on long prompts.
- CUDA graphs over trimmed layers 21 to 39 cut prefill compute 30 to 40%, dropping TTFT nearly 70% at 100K throughput.
- Integrated DeepSeek kernels: MegaAttention with NVFP4 KV (45% smaller), Mega-mHC, Mega-Gate, DeepSelect.
- Sparse MQA logits deliver 14 to 23x per-layer speedups at 512K context on NVIDIA GB300.
- GSM8K and GPQA show no meaningful accuracy regression from bounded replay; enabled by default.
vLLM speeds up DeepSeek V4.1-Flash serving
Three weeks after DeepSeek released V4.1-Flash, the vLLM team and Inferact reported serving optimizations that roughly doubled low-concurrency performance and delivered more than a 5× gain under heavy agent workloads. The work combines bounded sliding-window replay, broader CUDA graph coverage, compressed KV storage, and fused GPU kernels, as detailed in a technical write-up.
DeepSeek-V4.1-Flash is a 552-billion-parameter causal encoder-decoder designed for long-running agents. Its mixture-of-experts architecture activates 16 billion parameters per generated token during decode and 8 billion per prompt token during prefill. Compressed Sparse Attention 2, FP4 KV storage, and KV sharing between layers hold the global cache to about 890 bytes per token. The KV cache stores attention state from earlier tokens, so reducing its size directly increases the number and length of requests a server can handle.
Replay 128 tokens, skip half the stack
The model maintains two forms of attention state: a compressed global KV cache shared across layers and an FP8 sliding-window cache containing the latest 128 positions for each layer. Prefix caching, which reuses common prompt beginnings across requests, would need to save that sliding-window state at every possible cache boundary. According to vLLM, doing so requires more than 10× the storage of the global cache.
Standard prefill also sends every prompt token through layers 21 to 39, although decode consults only the latest 128 positions from those layers. vLLM’s bounded-replay path processes layers 0 through 20 across the full prompt to build the global state, then runs layers 21 through 39 only on the request’s final 128 tokens. Long prompts consequently avoid almost half of the model for all earlier positions.
An exact reconstruction would replay roughly 40 × 128 token steps through the stack. Bounded replay clips the sliding window at the replay boundary, which makes its output numerically different from full recomputation. Across tests including GSM8K and GPQA, all reported accuracy gaps remained within roughly 1.5 standard errors. The optimization is enabled by default and controlled with --[no-]swa-bounded-replay.
CUDA graphs recover launch overhead
Processing only 128 tokens through the upper layers leaves little GPU work per operation, allowing kernel-launch overhead to consume more of the runtime. CUDA graphs reduce that overhead by recording a sequence of GPU operations and replaying it without launching every kernel separately from the CPU.
vLLM uses piecewise graph capture for the two execution shapes: layers 0 through 20 run on the full batch, while layers 21 through 39 run on the trimmed batch. Bounded replay with graph capture reduced prefill computation time by 30% to 40%. Prompt length still matters: around 1,000 tokens, eager replay without graph capture ran as much as 12% slower than the original path.
Six kernel changes compound the gains
DeepSeek supplied kernels with the model, while vLLM integrated them and added further fusion and scheduling work. Kernel fusion combines operations into fewer GPU launches, reducing memory traffic and CPU scheduling costs.
| Optimization | Implementation | Reported result |
|---|---|---|
| MegaAttention with NVFP4 KV | Combines query rotary position encoding, sparse attention, inverse rotation, and FP8 casting in one launch. It reads KV data stored in NVIDIA’s 4-bit NVFP4 format. | 45% less KV storage than the previous FP8 cache and 1.45× kernel efficiency. |
| Sparse MQA logits | Uses DeepGEMM to score only the candidate positions considered by later indexer layers, which select the top results from as many as 16,000 candidates. | 1.2× faster at 8,000 tokens and 14× to 23× faster per layer at 512,000 tokens on NVIDIA GB300. |
| Mega-mHC | Fuses the post, delayed-pre, and RMSNorm stages into one kernel. | 1.14× to 1.51× faster than the previous TileLang path on GB200. |
| Mega-Gate | Combines the mixture-of-experts gate matrix multiplication, scoring, bias, and top-k selection. | 1.18× to 1.31× kernel speedups at medium batch sizes. |
| Fused WO-A | Uses a CuTe-DSL kernel on Blackwell GPUs to collapse three operations into one launch. | Roughly 6% to 7% lower inter-token latency at low concurrency. |
| mHC stream overlap | Calculates the next block’s coefficients on a secondary CUDA stream while attention and feed-forward work runs on the main stream. | About 4% faster low-latency decode with four-way tensor parallelism. |
Topology follows the bottleneck
The published results use the AgentX benchmark, which represents agent workloads with long shared prefixes, short bursts of generated tokens, and many turns. Time to first token, or TTFT, measures the delay between submitting a request and receiving its first generated token.
| Serving goal | Configuration | Reason |
|---|---|---|
| Low latency | Four-way tensor parallelism with FlashInfer attention | Small-batch decode is constrained largely by memory bandwidth, so distributing model weights across four GPUs reduces the work assigned to each device. |
| High throughput | DEP2 with data-parallel attention | Tensor parallelism would duplicate the model’s shared KV latent across GPUs, consuming memory that could serve additional requests. |
At the chart’s point near 100,000 throughput, the optimized configuration reduced TTFT by almost 70% relative to the initial serving stack.
Deployment trade-offs remain measurable
Agent services place sustained pressure on prefix caching, KV capacity, and launch latency because they repeatedly reuse long context while generating short responses. Bounded replay reduces upper-layer prefill work, compressed KV formats increase cache capacity, and fused kernels lower overhead during decode.
Deployments on Hopper and Blackwell hardware can obtain the main replay optimization by using a vLLM build that includes these integrations and retaining the default bounded-replay setting. Results will vary with prompt length, concurrency, GPU topology, and kernel support; several listed kernels target Blackwell specifically. Workloads that require bit-exact reproduction should disable bounded replay, and teams should validate application-level quality before rollout. Integration status and remaining work are tracked in issue #57448.