NVIDIA Ships Qwen2.5-VL Optimized for 3.6x Smaller Blackwell GPU Inference

NVIDIA published a 4-bit NVFP4 build of Qwen2.5-VL-7B-Instruct that runs on Blackwell Tensor Cores through TensorRT-LLM, cutting memory roughly 3.5x.

·
·
NVIDIA Ships Qwen2.5-VL Optimized for 3.6x Smaller Blackwell GPU InferencePRO
Read2 min
TypeModel
TopicImage · Gpus
  • NVIDIA released an NVFP4-quantized Qwen2.5-VL-7B-Instruct for Blackwell GPUs via TensorRT-LLM.
  • NVFP4 uses E2M1 4-bit values with FP8 E4M3 per-16-block scale plus per-tensor FP32 scale.
  • NVIDIA claims 2-3x throughput over FP8 and 3.5x less memory than BF16 with minimal accuracy loss.
  • Only language-model linear layers are quantized; 32K context and multimodal input preserved.
  • Calibrated on cnn_dailymail using the open-source TensorRT Model Optimizer v0.35.0.
  • Runs on DGX Spark today, B200 coming soon; EU deployment excluded under NVIDIA Open Model License.

NVIDIA packages Qwen2.5-VL 7B for native FP4 inference on Blackwell

NVIDIA has released a pre-quantized build of Alibaba’s Qwen2.5-VL-7B-Instruct for Blackwell GPUs. The model repository packages the vision-language model in NVIDIA’s NVFP4 format and targets deployment through TensorRT-LLM.

Blackwell Tensor Cores can execute NVFP4 natively, giving the checkpoint a path to lower memory use and higher inference throughput than BF16 or FP8 builds. Those gains apply mainly to the language model’s quantized linear layers; the vision encoder and several supporting tensors remain at higher precision.

NVFP4 squeezes weights into 16-value blocks

NVFP4 stores each quantized value in an E2M1 layout: one sign bit, two exponent bits, and one mantissa bit. Because four bits alone provide limited range, the format adds two levels of scaling:

  • Each 16-value micro-block shares an FP8 E4M3 scale.
  • Each tensor also carries an FP32 scale.
  • The FP8 block scales raise average storage to about 4.5 bits per value, with negligible additional cost from the per-tensor scale.

The storage ratio works out to roughly 3.6 times less than BF16 for quantized tensors. Blackwell also advertises twice the peak Tensor Core throughput for FP4 relative to FP8. End-to-end results depend on the vision encoder, KV-cache traffic, request batching, kernel availability, and the share of execution that remains at higher precision.

The vision stack stays at higher precision

NVIDIA produced the checkpoint from the Qwen base model with post-training quantization in Model Optimizer. The release has the following configuration:

Component Release detail
Quantized layers Linear operators in the language model’s transformer blocks
Higher-precision layers

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads