vLLM 0.31 Speeds Up DeepSeek and Slashes Engine Restart Times

The latest vLLM release lands DeepSeek-V4.1-Flash kernels, hot-restart weight caching, Model Runner V2 speculative decoding, and large-scale MoE serving upgrades.

·
·
·
vLLM 0.31 Speeds Up DeepSeek and Slashes Engine Restart Times
  • vLLM v0.31.0 ships 717 commits from 307 contributors with major DeepSeek-V4.1-Flash kernel upgrades.
  • FlashMLA mega attention with NVFP4 compressed KV cache is now the SM100 default.
  • New vllm preload CLI keeps post-quantized weights in GPU memory across engine restarts.
  • Model Runner V2 gains draft-model speculative decoding, LiLiCorr drafter, and async DFlash scheduling.
  • MoonEP all-to-all backend, DeepEPv2 with sequence parallelism, and NCCL M2N weight transfer for RL arrive.
  • Breaking: per-request mm_processor_kwargs rejected by default, tokenizer_mode=slow removed, enforce-eager now disables JIT warmup.

vLLM 0.31 speeds DeepSeek, keeps weights warm, and expands speculative decoding

vLLM 0.31.0 combines 717 commits from 307 contributors, including 96 first-time contributors. The release accelerates DeepSeek-V4.1-Flash on NVIDIA SM100 GPUs, adds a daemon that retains quantized weights during engine restarts, and brings draft-model speculative decoding to Model Runner V2.

For inference operators, the changes target token latency, GPU memory use, restart time, multi-node mixture-of-experts communication, and scheduler stability. Several renamed or removed options can also break existing launch configurations.

DeepSeek’s SM100 path gets fused

On NVIDIA SM100 Blackwell GPUs, vLLM now defaults to FlashMLA mega attention with DeepSeek-V4.1-Flash’s NVFP4 KV cache, which stores attention keys and values in a compressed 4-bit format. The indexer uses DeepGEMM’s sparse multi-query-attention logits kernel, while Mega-Gate combines the gate matrix multiplication and expert selection into one kernel launch. vLLM also folds several tensor-parallel collectives into adjacent kernels, including all-reduce with mHC input preparation and mixture-of-experts finalization. These changes reduce kernel-launch overhead and data movement.

Engram’s wkv tensors are now sharded across tensor-parallel ranks. Co-located data-parallel replicas also share Engram host lookup tables by default, removing per-process copies and increasing memory headroom as replica counts grow.

Warm weights shorten engine restarts

The new vllm preload command starts a weight-cache daemon that retains post-quantization weights in GPU memory while an engine process restarts. The daemon now supports data parallelism and multi-token-prediction draft models.

  • A /health endpoint exposes daemon health to orchestration systems.
  • A readiness wait lets deployment scripts block until cached weights are available.
  • Data-parallel deployments can reuse cached weights across engine restarts.
  • MTP draft-model weights can remain resident with the primary model.

Experimental snapshots add a second recovery path. The vllm snapshot create and vllm snapshot restore commands use Checkpoint/Restore in Userspace, or CRIU, to capture and restore a fully initialized engine. Snapshot support currently requires tensor-parallel size 1.

Keeping the cache daemon alive can remove disk reads, quantization, and part of engine warm-up from deployment and crash-recovery flows. A daemon failure, GPU reset, or host reboot still evicts the cached weights.

Draft-model decoding reaches Model Runner V2

Speculative decoding uses a smaller draft model to propose several tokens, then asks the primary model to verify them in batches. Model Runner V2 now supports this draft-model workflow, along with custom logits processors that modify token scores before sampling.

Additional Model Runner V2 work includes the LiLiCorr drafter, asynchronous DFlash scheduling with context keys and values precomputed in the draft CUDA graph, DSpark adaptive verification for Gemma4, and variable-length decoding for Kimi-K3.

Model Runner V2 profiling now counts mixture-of-experts allocations when estimating memory use. That correction prevents configurations from appearing viable during profiling and then exhausting GPU memory on WideEP deployments.

New backends target multi-node MoE bottlenecks

MoonEP adds a balanced expert-parallel all-to-all backend, selected with --all2all-backend moonep. All-to-all communication routes tokens among experts distributed across workers, often making network traffic a limiting factor in large mixture-of-experts deployments. Prefill context parallelism, which splits prompt processing across workers, can now operate with data parallelism.

DeepEPv2 adds sequence parallelism and expert-parallel load balancing, including overlap with shared-expert computation. The release also expands support for clusters that divide experts and token sequences across multiple nodes.

Reinforcement-learning systems gain a sharding-aware NCCL M2N weight-transfer backend. Training workers can send updated parameter shards directly to inference workers without rebuilding a complete model on each recipient. KV-cache offloading also detects backpressure, allowing the runtime to respond when external storage cannot keep pace.

Scheduler changes protect work in flight

  • --max-num-active-seqs limits running requests separately from the configured maximum number of sequences.
  • --long-prefill-token-threshold adapts to the number of prefills waiting in the queue, allowing a lone long prefill to proceed without unnecessary chunking.
  • The waiting queue prioritizes requests that already hold KV-cache blocks.
  • A deadlock involving the KV connector, MTP, and KV-cache pressure has been fixed.

Prioritizing partially processed requests preserves the KV-cache capacity already assigned to them. Under sustained load, this reduces starvation and avoids retaining cache blocks for requests that cannot resume.

Defaults that can break a rollout

Change Migration action
Per-request mm_processor_kwargs and media_io_kwargs are rejected by default. Set --trust-request-mm-kwargs only when clients are allowed to override multimodal processing and media I/O settings.
tokenizer_mode="slow" has been removed. Remove the value from configuration and test tokenizer selection under a supported mode.
--enable-mamba-fine-grained-prefix-cache has been renamed. Use --enable-mamba-shared-prefix-checkpoint.
The AllSpark INT8 W8A16 backend has been removed. Select another supported quantization backend before upgrading.
--enforce-eager now disables JIT kernel warm-up. Recheck startup and steady-state performance for eager-mode deployments.
XPU graphs are enabled by default. Revalidate graph capture, memory use, and execution behavior on Intel XPU systems.

CUDA 13 becomes the default build

PyPI now serves the CUDA 13.0 wheel by default. Pinning the package and container tag installs this release explicitly:

code
python -m pip install vllm==0.31.0
docker pull vllm/vllm-openai:v0.31.0

ROCm, XPU, and CPU builds are available alongside the CUDA package. A tagged CUDA 12.9 container remains available for deployments that have not moved to CUDA 13. The project is distributed under the Apache 2.0 license through the GitHub repository.

Upgrade targets and test plan

  • DeepSeek-V4.1-Flash deployments on SM100 Blackwell GPUs can use the new fused and compressed execution path.
  • Services with slow engine restarts can use the preload daemon to retain quantized weights.
  • Model Runner V2 users can adopt draft-model speculative decoding and custom logits processors.
  • Multi-node mixture-of-experts clusters can evaluate MoonEP, DeepEPv2, and the new parallelism combinations.

Production validation should cover launch flags, GPU architecture, prompt lengths, concurrency, quantization, and restart behavior. HiSparse and MTP received substantial changes in this cycle, while initialized-engine snapshots remain experimental and limited to tensor-parallel size 1.

The release touches nearly every major subsystem, so benchmark results from representative traffic provide the clearest upgrade signal. Staged rollouts can isolate scheduler, memory, and kernel regressions before the new version reaches the full fleet.

Trending
  • No trending articles

Comments

avatar

Next Reads