Stanford's ScholarCatalyst Reveals AI Misses 52% of Key Research Papers
A new benchmark asks AI to find the handful of prior papers that actually inspire new research, and today's retrieval agents fall short.
- ScholarCatalyst benchmarks whether AI can retrieve prior papers that actually inspired new research, not just topically similar ones.
- 184 lead authors across 207 CS projects labeled 894 queries with gold catalyst papers and rationales.
- Best system hits only 48% Recall@20 on a 191k-paper corpus restricted to pre-project literature.
- Tool-calling agents score 0.42 R@20, worse than the 0.48 embedding retriever they wrap.
- Even Claude Fable 5.1 reaches just 0.51 R@20; retrieval coverage, not reasoning, is the bottleneck.
- Code, data, and three agent baselines are open on GitHub and Hugging Face.
ScholarCatalyst exposes a recall gap in scientific search
ScholarCatalyst evaluates whether AI retrieval systems can find prior papers that materially shaped a research project while it was still taking form. In the main baseline comparison, the strongest embedding retriever surfaced 48% of the author-identified catalyst papers within its top 20 results.
Stanford’s IRIS Lab led the work with researchers from Seoul National University, the University of Washington, Carnegie Mellon University, the Allen Institute for AI, and MIT. The team collected relevance judgments from the authors of completed research instead of deriving them from citation graphs or topic labels.
Citation graphs miss the catalyst
Most scientific retrieval benchmarks treat citations, keywords, and topical similarity as evidence of relevance. Those signals identify neighboring work, but ScholarCatalyst targets papers whose ideas could materially advance a specific research question. The authors found that catalyst papers were no more semantically similar to a query than related papers the researchers had rejected.
ScholarCatalyst reconstructs the period before each project reached its key findings and evaluates retrieval under a temporal cutoff:
- Input: An early-stage research question validated by the researchers who pursued it.
- Search space: Literature published before the project’s key results emerged.
- Target: Prior papers that did, or plausibly could, have advanced the work.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.