UT Austin Reveals Why Efficient AI Models Fail at 10M-Token Reasoning
A new benchmark shows long-context LLMs handle retrieval fine but crumble on tasks like finding contradictions, breaking common architecture assumptions.
- New paper defines Corpus Task Complexity, grading tasks by how difficulty scales with corpus size.
- Most existing long-context benchmarks are low-CTC (linear); contradiction-finding and similar tasks are quadratic or worse.
- CTC-Bench packages 22 tasks, including 10 new high-CTC ones, evaluating up to 10M+ tokens.
- Block-sparse attention matches full attention on easy tasks but degrades sharply on high-CTC ones.
- Hybrid architectures like OLMo-3-Hybrid show similar growing gaps versus full attention on hard tasks.
- Short-to-long context generalization is much weaker for high-CTC tasks; code and data on arXiv.
Hard corpus tasks expose limits in efficient long-context models
Researchers at UT Austin, Carnegie Mellon University, and UC Berkeley have introduced a framework for measuring how reasoning difficulty grows with corpus size. Their paper, No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow, reports that sparse and hybrid attention models lose ground to full attention on tasks requiring comparisons across many documents.
Long-context evaluations commonly emphasize retrieval, single-hop question answering, and needle-in-a-haystack tests. Those workloads support claims about million-token context windows, but they reveal less about tasks that require models to connect, compare, or reconcile information throughout a corpus.
CTC tracks how corpus work grows
The paper defines Corpus Task Complexity (CTC) as the asymptotic number of oracle calls required to solve a task over a corpus of size N. The oracle is a conceptual tool, such as a language-model judge, that answers a bounded-context question in constant time relative to the full corpus. CTC measures how often an algorithm must invoke that oracle as the corpus expands.
- Linear complexity, O(N): A fact lookup may require scanning each document once. Work grows in proportion to corpus size.
- Quadratic complexity, O(N²): Exhaustive contradiction detection may require comparing every pair of claims.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.