LoopCD Cuts AI Reasoning Costs by 48% Without Changing Model Weights

A new training-free decoding trick for looped transformers contrasts early and late recurrent passes, boosting accuracy and cutting compute nearly in half.

·
·
·
LoopCD Cuts AI Reasoning Costs by 48% Without Changing Model WeightsPRO
  • Researchers introduce LoopCD, a training-free contrastive decoding method for looped transformers.
  • Contrasts final loop against an early loop to re-rank tokens, no auxiliary model needed.
  • Boosts Ouro-2.6B on AIME 2024 pass@1 from 61.88% to 73.33%.
  • Lifts Huginn HumanEval pass@1 from 22.56% to 31.71% with zero output overhead.
  • Enables halving recurrent loops, cutting forward FLOPs 22.5% to 48.2%.
  • Tested on Ouro, Huginn, Parcae, and Looped-Qwen3; first iteration is the best weak reference.

LoopCD Reuses Early Loops to Improve Transformer Decoding

Looped transformers reuse a shared block across recurrent iterations, increasing computational depth without adding a new set of weights at every step. Weihao Liu, Ruixiang Zhang, and collaborators propose LoopCD, a training-free decoding method that turns those iterations into a built-in contrastive signal. The authors, who completed part of the work during an Apple internship, report higher reasoning accuracy and forward-FLOP reductions of 22.5% to 48.2% when using fewer loops.

During ordinary inference, each recurrent iteration produces a hidden state that the model could map to next-token scores, or logits. Standard decoding uses the final state and discards the earlier predictions. LoopCD instead treats the first iteration as a weak predictor and the final iteration as a strong predictor, eliminating the auxiliary model that conventional contrastive decoding often requires.

One recurrent block, two predictors

LoopCD adjusts token scores by emphasizing the difference between the final and first recurrent states. Tokens that remain strong after additional computation receive more weight, while candidates favored mainly by the early iteration lose ground.

Variant Where contrast happens Additional output-stack calls Main trade-off
LoopCD-Logits Vocabulary logits One per generated token Applies contrast directly to token scores
LoopCD-Hidden Hidden states None Adds state storage and a vector operation

Both variants follow the same schematic update, with alpha controlling contrast strength:

code
Logit contrast:  z_cd = z_final + alpha * (z_final - z_first)
Hidden contrast: h_cd = h_final + alpha * (h_final - h_first)

In the logit variant, the model runs its output stack on both states before combining their vocabulary scores. In the hidden-state variant, it combines the states first and sends the result through the coda and language-modeling head once. The coda comprises any post-loop layers, while the language-modeling head maps the resulting vector to one score per vocabulary token.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads