Amazon's ALoDLM Beats Qwen3-8B With 2.7x Faster AI Text Generation

Amazon AGI's new diffusion language model spends more compute on hard tokens and less on easy ones, beating same-size autoregressive baselines across 11 benchmarks.

·
·
·
Amazon's ALoDLM Beats Qwen3-8B With 2.7x Faster AI Text GenerationPRO
  • Amazon AGI released ALoDLM, a diffusion language model with token-adaptive recurrent depth at each denoising step.
  • Easy tokens commit early as discrete context; hard tokens loop again, refining their latent states.
  • Trained at 1.7B and 8B scales; both beat Qwen3 AR baselines on average across 11 benchmarks.
  • ALoDLM-8B hits ~2.7x the throughput of vLLM-served Qwen3-8B at matched GSM8K accuracy.
  • Per-token exit schedules learned as latent variables via a conditional NELBO, no difficulty labels needed.
  • Weights and inference code open on Hugging Face and GitHub.

ALoDLM gives difficult tokens more compute

Researchers from Amazon AGI, the University of Illinois Chicago, and Korea University have released ALoDLM, a diffusion language model that adjusts computation for each token during generation. The release includes 1.7B and 8B parameter checkpoints, training code, and an optimized inference stack.

Standard diffusion language models predict multiple masked positions during each denoising step, creating more opportunities for parallel execution than left-to-right autoregressive models such as Qwen and Llama. Their quality has generally trailed comparable autoregressive models because every unresolved token receives the same processing depth, regardless of its difficulty.

According to the authors’ evaluation, ALoDLM outperforms all diffusion and autoregressive baselines they tested at both model sizes. On GSM8K, ALoDLM-8B delivers about 2.7 times the throughput of Qwen3-8B served through vLLM at comparable accuracy.

Fixed depth spends compute blindly

A diffusion language model begins with masked tokens and repeatedly fills or refines multiple positions. Within the same sequence, some predictions require little context, while others depend on several reasoning steps.

Function words often become predictable after nearby words appear. A number inside a multistep arithmetic solution may depend on calculations elsewhere in the sequence. Applying an identical transformer depth to both positions spends excess compute on the first and may leave the second underprocessed.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads