NVIDIA's Sigma Pushes Continuous Text Diffusion to 8 Billion Parameters
NVIDIA researchers unveil Sigma, a 3B/8B continuous diffusion language model that matches discrete rivals on math and code while offering smoother inference controls.
- NVIDIA-led team releases Sigma, the first 3B/8B continuous diffusion language model trained via likelihood.
- Jointly denoises Gaussian-corrupted token embeddings and learns the embedding geometry (constrained to a sphere).
- Warm-starts from pretrained AR weights to cut compute; adds auxiliary AR loss for stability.
- Classifier-free guidance plus score temperature identified as decisive inference knobs for reasoning and code.
- Competitive with masked dLMs and AR on GSM8K, HumanEval, MBPP, MATH-500 and AIME after SFT.
- Parallel Decoding Distillation (PDD) preserves quality at low step counts, enabling cheaper inference.
Sigma scales continuous diffusion language models to 8 billion parameters
Diffusion image models generate samples by repeatedly denoising continuous pixels or latent vectors in parallel. Text diffusion systems usually operate on discrete vocabulary tokens, revealing masked positions over several steps. In the Sigma paper, an NVIDIA-led team instead denoises continuous token embeddings. The authors describe Sigma as the first likelihood-trained continuous diffusion language model at 3-billion and 8-billion parameter scale.
| Model family | Working state | Generation process |
|---|---|---|
| Autoregressive | Token prefix | Predicts the next token from previous tokens. |
| Masked diffusion | Masked and visible tokens | Predicts multiple masked positions during each decoding step. |
| Continuous diffusion | Noisy token embeddings | Refines continuous vectors before mapping them back to vocabulary tokens. |
Token masks make steering awkward
Masked diffusion systems such as LLaDA and Dream 7B can update many positions within a block, reducing the sequential work required by autoregressive decoding. Each position still jumps among a mask and tens of thousands of vocabulary entries. Those categorical transitions provide no smooth path through which a sampler can adjust its direction.
Sigma places each token in a learned continuous space whose dimensions are far fewer than a one-hot vocabulary representation. Continuous vectors define smooth sampling trajectories, represented by an ordinary differential equation for deterministic sampling or a stochastic differential equation that injects noise. That structure supports guidance and temperature controls already used in image diffusion.
Inside Sigma’s continuous state
Sigma replaces masked-token corruption with Gaussian noise applied to token embeddings. The model learns to reverse that corruption while jointly learning the embedding geometry. After denoising, a token-prediction head converts each vector into a probability distribution over the vocabulary.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.