Stanford's Priced Guidance Framework Measures How Hard AI Finds Research Ideas
Stanford researchers introduce a compression-based framework that quantifies how close language models come to generating future research ideas, measured in bits.
- Stanford's Priced Guidance measures how many bits are needed to steer an LLM to a future paper's core idea.
- Framework proves K bits of guidance implies unguided generation probability of at least 2^(-K).
- Fable 5.1 hits median 69.9 bits per paper, versus 5,712 bits for gzip on the same summaries.
- Ensemble of Fable 5.1, Opus 5, and Astra drops median to 55.8 bits, an 18,000x probability improvement.
- Tested on 87 recent Hugging Face Daily Papers; 61% of Opus 5's bits go to keyword questions.
- Code and agent trajectories at github.com/WhenWen/priced-guidance.
Priced Guidance Puts a Bit Cost on AI-Generated Research Ideas
Stanford researchers Kaiyue Wen, Tengyu Ma, and Percy Liang have proposed a way to evaluate an unusually rare capability: whether a language model could generate the core idea of a research paper without seeing it. Direct sampling breaks down when a successful idea may appear once in billions or trillions of attempts. Their Priced Guidance preprint measures how many bits of help the model needs to reach the paper’s idea.
By connecting prediction with compression, the framework treats likely ideas as cheap to specify and unexpected ideas as expensive. The resulting bit cost provides a mathematically derived lower bound on the probability that the model would produce the same idea without guidance.
A rare event becomes a bit budget
- The generator asks a question. A language model proposes an adaptive multiple-choice question about the unknown paper and assigns a probability to each answer.
- The guide selects an answer. A second model, which has read the target paper, chooses the option that best describes it.
- The framework charges for the answer. Answers the generator considered likely cost fewer bits; surprising answers cost more.
- A judge evaluates the result. An independent model checks how closely the reconstructed proposal matches the target paper.
The judge uses two levels of success:
| Level | Criterion |
|---|---|
| Directional | The proposal identifies the same broad research direction. |
| Essence | The proposal recovers the central research object and defining mechanism. |
For an answer assigned probability p, the price is -log2(p) bits. An answer with probability 0.5 costs one bit, while an answer with probability 0.01 costs about 6.64 bits. The framework adds these prices across the full question sequence.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.