Simplex Diffusion Models Beat Discrete Rivals at Math With 16x Fewer Steps
A new discrete diffusion framework operates directly on the probability simplex, carrying uncertainty between denoising steps and beating masked diffusion on code and math benchmarks.
- New Simplex Diffusion Models keep denoising on the probability simplex, avoiding information collapse in discrete diffusion.
- Dirichlet forward process with decoupled mean and concentration yields closed-form reverse transitions, no ODE integration required.
- DDIM-like sampler with a tunable churn parameter, trained with plain cross-entropy loss.
- Hits 17.0 GenPPL on OpenWebText in 64 steps, near real validation data.
- 49.0% on TinyGSM at T=0.1 without self-conditioning, beating masked and uniform diffusion baselines.
- Distilled to 8 steps, solves 32.1% of GSM8K, versus 21.4% for distilled discrete diffusion at 128 steps.
Simplex Diffusion Models keep uncertainty between steps
Many discrete diffusion samplers discard useful information during generation: each denoising step produces a categorical distribution, collapses it into one token, and passes that hard choice to the next step. A new paper, Simplex Diffusion Models, proposes keeping the full probability vector throughout the trajectory. The approach preserves token uncertainty, supports closed-form sampling updates, and improves reported results on text, math, and code benchmarks.
Where hard tokens lose information
Continuous diffusion refines a real-valued state over many small updates. Standard discrete diffusion commonly samples or selects a categorical token after each update, then re-encodes that token for the next network call. This conversion erases the relative probabilities assigned to alternative tokens.
Self-conditioning partly restores that information by feeding the previous soft prediction back into the model. Predictor-corrector samplers add further refinement calls. Both techniques can improve quality, although they increase model complexity or inference cost.
Dirichlet Flow Matching previously relaxed categorical data onto the probability simplex, the set of nonnegative vectors whose components sum to one. Its sampler requires numerical integration of an ordinary differential equation, and linear flow matching near categorical endpoints can produce discontinuous vector fields.
A smoother path through the simplex
An SDM's forward process samples simplex-valued states from a Dirichlet distribution. Its mean moves from the clean one-hot token toward a uniform prior, while an independently controlled concentration parameter sets the distribution's spread. High concentration keeps samples near the mean; low concentration produces more variable, corner-seeking samples.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.