Red Hat Shrinks Nemotron 3.5 Lightning's 30B Agent Model by Half With FP8
Red Hat AI shipped an FP8 build of NVIDIA's hybrid Mamba-MoE Lightning model, halving memory while keeping the 1M-token agent workhorse on a single GPU.
- Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B.
- Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.
- Base model is a hybrid Mamba-2 + MoE + Attention architecture with 3B active parameters.
- Supports up to 1M token context, switchable reasoning, tool calling, and six languages.
- Scores 83.4 on PinchBench agent tasks but lags on coding and hard reasoning benchmarks.
- Deployable via vLLM, SGLang, or Docker on a single H100 or DGX Spark.
Red Hat cuts Nemotron 3.5 Lightning’s weight footprint with FP8
Red Hat AI has released an FP8-quantized checkpoint of NVIDIA’s Nemotron model. Converting selected weights and activations from BF16 to FP8 roughly halves their storage, making the 30-billion-parameter agent model easier to deploy on a single accelerator.
Hugging Face displayed more than 780,000 downloads for the checkpoint at the time of publication. That counter measures file downloads rather than unique users or production deployments, but it indicates substantial early interest in a lower-memory version.
Half-size weights, qualified gains
Red Hat produced the checkpoint with LLM Compressor, applying static per-tensor FP8 quantization to weights and activations in supported linear operators. FP8 stores each quantized value in eight bits, compared with 16 bits for BF16.
Several precision-sensitive components remain unquantized: conv1d layers, embeddings, latent projections, mixture-of-experts gates, multi-token prediction layers, the final normalization layer, and the language-model head. Calibration used 512 UltraChat samples at 2,048 tokens each.
The resulting savings apply primarily to quantized tensors. Total runtime memory also includes unquantized parameters, activations, Mamba state, attention key-value caches, and vLLM workspace allocations, so overall GPU memory use falls by less than 50% in many serving configurations.
Hopper and Blackwell GPUs can execute FP8 matrix multiplication through dedicated Tensor Cores. End-to-end throughput depends on batch size, sequence length, memory bandwidth, expert routing, and the share of execution handled by unquantized operators. FP8 therefore raises the performance ceiling without guaranteeing a twofold increase in request throughput.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.