SmolLM3 Retrieves Perfectly at 64x Its Training Length by Moving Dropout

A new paper argues the right dropout placement, not just positional encoding, lets transformers extrapolate up to 64x past their training context length.

·
·
·
SmolLM3 Retrieves Perfectly at 64x Its Training Length by Moving DropoutPRO
  • New paper shows regularization choices, not just positional encoding, drive length generalization in transformers.
  • Weight decay hurts extrapolation; dropout helps when moved from pre-norm to pre-linear placement.
  • Standard pre-norm dropout creates a variance mismatch with LayerNorm that worsens with sequence length.
  • Modified SmolLM3 extrapolates perfectly to 64x training context on Needle-in-a-Haystack with the fix.
  • Proposed Variance-Preserving Affine Dropout (VPAD) further shrinks the pre-activation variance gap.
  • Mamba2 also benefits, indicating the dropout fix is architecture-agnostic across transformers and SSMs.

Dropout placement changes long-context extrapolation

Length generalization measures whether a model trained on short sequences can handle much longer ones. Common remedies alter positional encoding or attention through NoPE, ALiBi, RoPE scaling, YaRN, NTK scaling, or sliding windows. Researchers at Instituto Superior Técnico and the University of Edinburgh now report that regularization choices can determine whether those architectures extrapolate successfully.

The arXiv paper identifies two effects across its experiments. Weight decay weakened extrapolation, while dropout improved it when moved to a different point inside each block. A modified SmolLM3 continued pretraining with the revised recipe achieved perfect Needle-in-a-Haystack retrieval at tested lengths through 64 times its pretraining context.

Why the mask’s location matters

In a typical pre-norm transformer block, normalized hidden states pass through attention or an MLP, dropout follows the output projection, and the result joins the residual stream. A simplified form is x_next = x + Dropout(F(Norm(x))). The next block receives that altered residual stream through LayerNorm, RMSNorm, or a related normalization layer.

During training, inverted dropout keeps each activation with probability q and scales survivors by 1/q. For a fixed activation x, the masked value z = mx/q preserves the expectation, E[z] = x, while increasing the second moment to E[z²] = x²/q. At inference, dropout disappears and the unmasked activations pass through unchanged.

The relevant train-test mismatch sits in the residual vectors presented to normalization. LayerNorm and RMSNorm compute statistics anew for each token, but stochastic masking still changes the vector’s variance, direction, and mixture with the residual branch. The paper links the accumulated mismatch across layers to degraded performance beyond the training length.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads