Exa Labs' ATLAS Exposes How AI Agents Fake Web Research Without Searching

Exa's new ATLAS benchmark exposes how saturated and memorized existing search evals are, with top agents still missing a third of answers.

·
·
·
Exa Labs' ATLAS Exposes How AI Agents Fake Web Research Without Searching
Read5 min
TypeNews
  • ATLAS is Exa's new 547-task benchmark for agentic web search, targeting deep and wide research workflows.
  • Only 9% of tasks are memorized by frontier models, vs 48 to 61% for BrowseComp, WideSearch, and DeepSearchQA.
  • Best system hits 0.66 row F1 at $8.92 and 18 minutes per task; no sub-$1 system clears 0.5.
  • Hiding top 7 of 10 search results halves ATLAS scores but barely dents other benchmarks.
  • Tasks built via a 6-stage pipeline using Codex, Claude Code, and a provider-neutral Exa/Brave/Perplexity search tool.
  • Grading costs about $2 for 10 full runs; tasks, gold tables, and grader releasing in coming weeks.

ATLAS targets memorization in search-agent benchmarks

Exa Labs has introduced ATLAS, a benchmark designed to test how well research agents retrieve broad, detailed answers from the web. Its 547 tasks emphasize exhaustive entity discovery and multi-step fact gathering, exposing differences among search APIs that older benchmarks may obscure.

Memorization blurs retrieval scores

Reported scores have reached 93% on BrowseComp and 95% on DeepSearchQA, leaving little room to distinguish search systems. Exa’s no-browsing tests found that at least one frontier model could reproduce answers for 48% to 61% of tasks from BrowseComp, WideSearch, and DeepSearchQA. The company classifies those tasks as memorized because the models completed them without live retrieval.

Exa reports a 9% memorization rate for ATLAS. Perplexity’s newer WANDR benchmark also addresses stale questions, though its LLM-based grading adds inference costs and makes results sensitive to the grader model and prompt. ATLAS uses a fixed answer key that reportedly costs about $2 to grade 10 complete benchmark runs.

Rows force breadth and depth

Each ATLAS task requires an agent to identify every entity that satisfies a set of conditions, then populate each entity’s row with two to 10 attributes. Many attributes require multi-hop research, meaning the agent must connect facts gathered through several searches or sources.

Seventy-four percent of tasks require at least 10 entities. Discovery draws from a median of four web domains per task, while discovery and enrichment together use a median of 18. A strong agent completed 48% of WideSearch tasks with citations from one site; the equivalent rate on ATLAS was 0.2%.

ATLAS reports three F1 variants, which balance precision against recall. Its headline metric is Row F1: a row counts as correct only when every cell is correct. This metric penalizes missing entities, incorrect additions, and partially completed rows.

Search quality shows up in the score

Exa’s reported ATLAS results
Measure Result
Best tested Row F1 0.66 on a 0-to-1 scale
Cost of the best setup $8.92 per task
Latency of the best setup 18 minutes per task
Systems costing under $1 per task All scored below 0.50 Row F1
Search API spread with a fixed agent setup 16 Row F1 points

A retrieval-degradation test hid seven of every 10 top search results from the agents. ATLAS scores fell by roughly half, while BrowseComp, WideSearch, and DeepSearchQA declined by 11 to 19 points. The larger drop indicates that ATLAS depends more heavily on the quality of retrieved results.

Benchmark scores after top search results were hidden
ATLAS showed a larger response to degraded retrieval than the older benchmarks in Exa’s tests.

With the model and agent harness held constant, switching among Exa, Perplexity, OpenAI, Brave, and Parallel produced the reported 16-point spread. Exa-backed configurations occupied the reported cost-performance Pareto frontier, meaning no measured alternative was both cheaper and more accurate.

ATLAS Row F1 plotted against cost for several search APIs
Exa’s comparison plots Row F1 against task cost for different search backends.

Because Exa created the benchmark and sells one of the evaluated search APIs, the backend comparison requires independent replication. The underlying tasks, answer tables, and grader were not public when the results were announced.

Building the answer key

Exa says each task required a median of eight aggregate agent-hours and 1,200 searches. Multiple agents gathered answers, checked them against cited sources, and resolved disagreements through a six-stage pipeline:

  1. Topic seeding: Exa sampled anonymized aggregate search demand and distributed tasks across more than 300 topics.
  2. Entity discovery: Codex with GPT-6 Astra and Claude Code with Opus 5.5 searched through a tool that interleaved Exa, Brave, and Perplexity results while concealing provider labels.
  3. Row enrichment: Agents filled attribute cells with citations and verbatim supporting quotes.
  4. Stress testing: Twelve agent configurations attempted the resulting tasks.
  5. Blind audit: Codex and Claude Code reviewed the answers, with Gemini 3.1 Pro serving as a tie-breaker.
  6. Answer-key assembly: The final tables tracked ungraded cells at 7.1% and confirmed blank values at 3.8%.

A separate live-browser audit estimated that 0.9% of recorded gold values were incorrect, with a 95% confidence interval of 0.6% to 1.4%. That estimate addresses errors in recorded values; it does not by itself establish that every valid entity appears in the answer key.

Cheap grading, pending artifacts

A fixed answer key could make ATLAS practical for continuous evaluation. Exa regraded 7,658 answers without changing any system’s mean Row F1 by more than 0.001, indicating that repeated grading is stable under the reported setup.

At publication, Exa said it would release the tasks, gold tables, and grader within weeks. The company also plans to rebuild the benchmark from fresh search demand every three to six months, which should reduce exposure to memorization as models absorb public evaluation sets.

Teams comparing search providers can use the benchmark to isolate retrieval quality once the artifacts are available:

  • Fix the model, prompts, tool interface, and task budget.
  • Change only the search API between runs.
  • Track Row F1 alongside cost, latency, and task failures.
  • Repeat the comparison when Exa publishes a refreshed task set.

What the scores mean for products

Products that depend on exhaustive coverage, including structured research, lead generation, and compliance reviews, remain sensitive to search-backend quality and retrieval budgets. Exa’s results show substantial room for error across every tested setup, especially below $1 per task. Public artifacts and independent runs will determine whether ATLAS can reproduce those distinctions outside Exa’s evaluation environment.

Trending
  • No trending articles

Comments

avatar

Next Reads