vLLM Runs MiniMax M3 at 7.84x Faster on NVIDIA's Vera Rubin
vLLM lands day-0 support for NVIDIA Vera Rubin NVL72 with Rubin-tuned kernels and locality-aware MoE, hitting 7.8x the per-GPU throughput of GB200.
- vLLM adds day-0 support for NVIDIA Vera Rubin NVL72 via nightly cu134 containers.
- Delivers 7.8x per-GPU throughput over GB200 on SemiAnalysis AgentX with MiniMax M3.
- MLPerf Inference v6.1: 3.7x higher VLM throughput versus GB300 NVL72 on Qwen3-VL-235B.
- Rubin brings 5x NVFP4 FLOPS, 2.4x HBM bandwidth, 1.7x NVLink, and 2-4x softmax exponentials.
- Locality-aware MoE using CUDA 13.4 locality domains gives 1.2x speedup on MiniMax M3 MoE decode.
- Supports DeepSeek, Kimi, GLM, MiniMax via
vllm/vllm-openai:cu134-nightlycontainer today.
vLLM has added preview support for NVIDIA’s Vera Rubin NVL72 platform in nightly containers built with CUDA 13.4 and PyTorch 2.15. Working with Inferact, NVIDIA, and Red Hat, the project has run DeepSeek, Kimi, GLM, and MiniMax models on Rubin hardware. FlashInfer 0.7.0 supplies Rubin-tuned attention, GEMM matrix-multiplication, and mixture-of-experts (MoE) kernels.
Early results in the vLLM preview show substantial throughput gains over NVIDIA’s GB200 and GB300 systems. The comparisons cover two different workloads and remain preview measurements rather than independently reproduced production benchmarks.
Two workloads, large early gains
The AgentX results use MiniMax M3 to test agentic inference, where long contexts and repeated model calls place pressure on memory bandwidth and MoE execution. The MLPerf Inference v6.1 result uses the multimodal Qwen3-VL-235B-A22B model, with NVIDIA Dynamo routing requests to vLLM workers.
| Benchmark | Model and stack | Baseline | Condition | Reported gain |
|---|---|---|---|---|
| AgentX | MiniMax M3 on vLLM | NVIDIA GB200 | Matched interactivity | Up to 7.84× throughput per GPU |
| AgentX | MiniMax M3 on vLLM | NVIDIA GB200 | 150 tokens per second | 5.18× throughput |
| MLPerf Inference v6.1 VLM | Qwen3-VL-235B-A22B with vLLM and Dynamo | GB300 NVL72 | Offline, server, and interactive scenarios | Up to 3.7× throughput |
The figures describe separate benchmark configurations, so they should be evaluated within their respective workloads. “Matched interactivity” holds the benchmark’s responsiveness target constant while measuring how much aggregate traffic each system can serve.
More bandwidth for the decode path
Vera Rubin NVL72 connects 72 Rubin GPUs as a rack-scale system. Rubin retains compatibility with much of Blackwell’s programming model while adding the sm107 compile target and extending the tcgen05 tensor-core instruction set. NVIDIA reports the following peak per-GPU improvements over GB200 NVL72:
- 5× more NVFP4 inference FLOPS
- 2.4× more HBM bandwidth through the move from HBM3e to HBM4
- Up to 1.7× more inter-GPU bandwidth through sixth-generation NVLink
- 2× FP32 and 4× BF16/FP16 exponential throughput
LLM inference typically starts with a compute-heavy prompt-processing phase and shifts into memory-bound token generation. HBM4 helps the decode phase by moving model weights and key-value cache data faster. Higher exponential throughput also accelerates softmax, an attention operation whose gains have historically trailed improvements in matrix multiplication.
MoE weights move closer to compute
CUDA 13.4 exposes locality domains that let software place memory and computation near each other within a GPU. Global memory access has been non-uniform on NVIDIA GPUs since Ampere, meaning a streaming multiprocessor can read some HBM regions faster than others. Locality domains give runtimes a way to exploit that topology directly.
For MoE decoding, vLLM divides the FC1 and FC2 expert-weight matrices into column-wise shards, assigns each shard to a locality domain, and uses CUDA Green Contexts to launch a kernel for each domain. Streaming multiprocessors then read nearby expert weights instead of repeatedly crossing the GPU’s internal memory fabric.
vLLM reports an average 1.2× speedup for MiniMax M3 MoE layers in small-token forward passes. The result held across TP2, TP4, EP2, and EP4 configurations, where TP denotes tensor parallelism, EP denotes expert parallelism, and the numeral indicates the parallel group size.
Pull the CUDA 13.4 nightly
Developers with Rubin access can pull the project’s CUDA 13.4 nightly image with the following command:
docker pull vllm/vllm-openai:cu134-nightlyThe main tuning flags select Rubin-capable kernels and communication paths:
--linear-backend flashinfer_cutedslselects the CuTe DSL NVFP4 dense GEMM backend. It is enabled by default for NVFP4 checkpoints.--moe-backend flashinfer_cutedslenables the CuTe DSL NVFP4 MoE kernel.--kv-cache-dtype fp8with--attention-backend FLASHINFERorFLASHINFER_MLAenables the TensorRT-LLM-generation FP8 attention path.--enable-expert-parallel --data-parallel-size N --all2all-backend deepep_low_latencyenables expert parallelism, DeepEP low-latency communication, and masked grouped GEMM for batched expert formats.
Current model coverage includes DeepSeek, Moonshot AI’s Kimi, Z.ai’s GLM, and MiniMax. Rubin can execute existing Blackwell kernels built for the forward-compatible sm100f target, which provides broad initial coverage while dedicated sm107 kernels continue to arrive. Because nightly images are mutable, production evaluations should pin an image digest and record the vLLM, FlashInfer, CUDA, and driver versions.
Production math needs more data
Under ideal linear scaling, a 7.84× per-GPU throughput gain would reduce the GPU count for the same AgentX request volume to about 13% of the GB200 requirement. The 5.18× result would reduce it to about 19%. Actual capacity planning will depend on utilization, tail latency, model shape, context length, batching, power, rack cost, and whether an application resembles the measured workload.
- The results come from early hardware and a limited number of nodes.
- The collaborators published the measurements, and independent reproductions are not yet available.
- “Up to” figures identify the best reported configuration rather than a guaranteed gain across models and traffic patterns.
- The published comparisons do not provide enough cost and power data for a full total-cost-of-ownership analysis.
- CUDA locality-domain support remains under active design and development.
vLLM’s roadmap includes broader locality-domain support for MoE layers, an sm107 FlashInfer MegaMoE path, and fused mega-kernels that reduce kernel-launch and memory overhead on latency-sensitive requests. The project also plans Rubin-specific KDA, MLA, CSA, and HCA attention kernels for Kimi K3 and DeepSeek-V4.1-Flash. Those additions will determine how much performance the stack can extract beyond its current Blackwell-compatible execution paths.