NYU's HyperThink Cuts AI Reasoning Costs 10x Without Long Thinking Traces

A new COLM 2026 paper proposes replacing long chain-of-thought traces with a tiny, question-specific weight patch generated on the fly.

·
·
·
NYU's HyperThink Cuts AI Reasoning Costs 10x Without Long Thinking TracesPRO
  • New COLM 2026 paper HyperThink replaces chain-of-thought traces with query-specific weight updates.
  • A hypernetwork predicts bias offsets, under 0.02% of parameters, routed through a VQ codebook.
  • Approaches full thinking-mode accuracy with up to 3.9x fewer FLOPs on math benchmarks.
  • Brings Olmo-3-7B-Think near full thinking Pass@5 while answering up to 10x faster.
  • Codebook entries turn out interpretable, specializing by math subject without supervision.
  • Code will be open-sourced; co-led by Jack Lu and Donggyun Kim across NYU and KAIST labs.

HyperThink shifts LLM reasoning into per-query weight updates

Researchers from the NYU Agentic AI Lab and KAIST Vision and Learning Lab have proposed HyperThink, a system that compresses language-model reasoning into temporary, query-specific bias updates. A hypernetwork reads the question, generates a small patch for a frozen base model, and lets that model produce an answer without first decoding a long chain-of-thought trace. Jack Lu and Donggyun Kim co-led the work, which was accepted at COLM 2026.

Reasoning carries a sequential bill

Reasoning models often improve accuracy by generating thousands of intermediate tokens before returning an answer. Autoregressive decoding produces those tokens one at a time, so longer traces increase latency, compute use, and API cost even when users only need the final response.

HyperThink moves part of that computation into one pass through a small side network. The base model still decodes the answer normally, but it begins with query-specific bias offsets intended to encode information that would otherwise emerge during a longer reasoning trace.

A temporary patch for each query

Each HyperThink request follows four stages:

  1. A frozen text encoder converts the question into a representation.
  2. A bias encoder maps that representation to a set of bias tokens.
  3. A vector-quantized decoder matches those tokens to entries in a learned codebook and produces bias offsets.
  4. The serving system applies the offsets to selected later layers of the frozen LLM, then decodes the answer.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads