Exa Beats Google Scholar by 3x Finding Academic Papers Researchers Half-Remember
Exa launches semantic search over 350M scientific papers, outperforming Google Scholar with 86.4% recall at sub-400ms latency

- Exa launches semantic search over 350M+ scientific publications, queryable via a new
category="publication"API parameter. - Benchmarks show 86.4% recall vs Google Scholar's 28%, with sub-400ms latency -- more than 3x more accurate and faster.
- Two new retrieval benchmarks released publicly: known-item and tip-of-the-tongue retrieval for research papers.
- Ingestion pipeline uses OCR + document parsing to index PDFs, tables, appendices, and scanned documents, combined with author/citation metadata.
- Harvard Medical School professors and biopharma companies like Elicit are already using it for scientific AI agents.
- Available now through the Exa API; free tier includes 1,000 queries/month with no credit card required.
Exa launched a dedicated academic search product with benchmark numbers worth examining. The company built an index of roughly 350 million publications and added around 30 million authors to its people index, all queryable through its existing Search API. The goal: scientific search that retrieves papers the way researchers think about them, not the way databases catalog them.
Why academic search keeps failing researchers
Traditional tools like Google Scholar rely on exact keyword matching against titles, author names, and abstracts. That works when you already know what you're looking for. It breaks down when all you have is a vague memory of a result or a question like: "wasn't there a study from the 70s showing dietary cholesterol raised LDL in primates even when total cholesterol looked normal?"
Exa searches over the meaning of a query and the full contents of a paper, so you can retrieve a specific publication from factual clues or an incomplete recollection without translating your question into academic keywords. The shift is from keyword lookup to semantic retrieval.
The benchmark numbers
Exa created two benchmarks to measure this. The first tests known-item retrieval, where a searcher provides precise factual clues and the system must identify the exact paper. The second tests tip-of-the-tongue retrieval, where the query is vague, incomplete, or slightly wrong, closer to how someone recalls a paper they read months ago.

Results across both benchmarks:
| System | Recall | MRR | Mean Latency |
|---|---|---|---|
| Exa | 86.4% | 0.726 | 0.578s |
| Perplexity | 66.8% | 0.568 | 1.277s |
| Parallel Advanced | 50.0% | 0.312 | 3.118s |
| Parallel Turbo | 39.2% | 0.278 | 0.403s |
| Google Scholar | 28.0% | 0.179 | 1.098s |
MRR (Mean Reciprocal Rank) measures how high the correct paper appears in the result list; a score of 1.0 means it's always first. Exa led on both recall and MRR while returning results in under 600 milliseconds. Its recall is more than three times Google Scholar's, at faster speeds.
How the index was built
Research papers are difficult documents to search. Key information can sit in a PDF appendix, a scanned table, or a supplementary file. The same paper often appears at multiple URLs with inconsistent metadata. Exa's ingestion pipeline addresses this directly:
- An OCR and document-parsing pipeline converts research PDFs into clean, searchable text.
- That text is combined with structured metadata: authors, institutions, publication histories, citations, and collaborator graphs.
- At query time, Exa searches both its web index and the dedicated publication index, then merges and reranks results, giving researchers open-web coverage alongside a purpose-built scientific index.
Underlying infrastructure includes crawlers tracking hundreds of billions of URLs, custom retrieval and embedding models trained on Exa's own GPU cluster, and vector databases sized for the query volumes that AI agents generate.
One gap in the current benchmarks
Exa is developing a third benchmark it doesn't yet have results for: topical completeness. Rather than finding one specific paper, it measures how thoroughly a search system recovers the important body of work around a topic. Finding a single known paper is useful; understanding a field requires surfacing the surrounding literature too. Until that eval ships, the current benchmarks capture only part of what scientific agents actually need.
How to use it
Exa released a new publication search category, which replaces the older research_paper category. If you're already using the Exa API, the change is a single argument:
from exa_py import Exa
exa = Exa()
results = exa.search(
"papers proposing alternatives to attention for long-sequence modeling",
category="publication",
num_results=10
)You can also test it in the Exa dashboard without writing any code. Pricing follows Exa's standard API model: $7 per 1,000 requests with text and highlights included, $1 per 1,000 for summaries, and $12–$15 per 1,000 for Deep and Deep-Reasoning modes. A free tier covers 1,000 queries per month.
Who's using it
Harvard Medical School professors Omar Abudayyeh and Jonathan Gootenberg described Exa as "core to how we build deep thinking scientific agents, from biomedical research to medical analysis across complicated scientific domains." Exa already powers search for Cursor, Cognition, HubSpot, OpenRouter, Monday.com, and over 400,000 developers. Elicit, a research assistant built for scientists, has integrated Exa into its biopharma workflows.
Where this fits in Exa's broader strategy
Exa builds search infrastructure for AI agents and recently closed a $250 million Series C at a reported $2.2 billion valuation. The publications launch extends that thesis into one of the highest-value domains for AI retrieval: drug discovery, clinical research, and materials science all depend on finding the right papers quickly, and the volume of LLM-driven searches is growing fast enough that retrieval quality will directly affect downstream output quality.
For teams building research agents, the practical case is concrete. If your pipeline calls Google Scholar or a generic web search API for scientific context, switching to Exa's publication category is a one-argument change. A recall gap of 86% versus 28% is wide enough to surface in answer quality, not just in retrieval metrics.