BAAI's AREX Beats Models 10x Its Size on Deep Web Research
BAAI's AREX agent audits its own answers constraint by constraint, keeping verified pieces and researching only the gaps, hitting 82.5% on BrowseComp.
- BAAI released AREX, a recursively self-improving research agent with inner and outer loops.
- The outer loop audits answers constraint by constraint, keeping verified parts and re-researching unresolved ones.
- Two models: AREX-Turbo (4B dense) and AREX-Base (122B MoE, 10B active), both built on Qwen3.5.
- Scores 82.5 on BrowseComp, 82.0 on WideSearch-en, 85.4 GAIA, 89.9 DeepSearchQA.
- Harness plus context updating adds 22.9 points over linear search baseline on BrowseComp.
- Fine-tuned 4B Turbo beat untuned Qwen3.5-35B on five of six benchmarks.
Deep-research agents often waste tool calls by repeating failed searches, accepting partially supported answers, and feeding entire traces back into a crowded context window. A team at the Beijing Academy of Artificial Intelligence has released the AREX paper, which describes a family of agents that verifies and revises answers throughout the research process. The approach preserves useful findings while directing subsequent searches toward specific evidence gaps.
AREX builds on the premise that finding an answer with several independent requirements can be expensive, while checking a candidate against each requirement is more tractable. Every check updates a compact state containing supported claims, unresolved constraints, and the evidence gathered so far. A partially supported answer therefore becomes the starting point for a narrower research pass.
Two loops keep the search moving
AREX divides each research task into two nested loops:
- Inner research loop: The model follows a ReAct-style cycle, alternating between reasoning, tool calls, and observations as it searches the web, reads sources, and drafts a provisional answer.
- Outer improvement loop: The model audits that answer against each requirement in the original question, records unresolved claims, and generates targeted follow-up queries.
Long trajectories are managed through an autonomous context-update tool. It compresses the interaction history into an improvement state that retains verified evidence and open constraints. The agent chooses when to invoke the tool and which details survive, with no separate model performing the compression.
Stopping also becomes an explicit agent decision. AREX emits a confidence score from 0 to 100 when it considers an answer complete. A score above the configured threshold returns the answer; a lower score triggers another targeted pass or a new research pass seeded with the verified findings.
Training targets consequential decisions
BAAI provides two Qwen3.5-based implementations:
| Model | Architecture | Total parameters | Active parameters |
|---|---|---|---|
| AREX-Turbo | Dense | 4B | 4B |
| AREX-Base | Mixture of Experts | 122B | 10B |
A Mixture-of-Experts model routes each token through a subset of its parameters, which explains the gap between Base's total and active counts. BAAI publishes the model weights and source code.
Long agent trajectories contain many routine actions and relatively few decisions that change the outcome. The authors use key-step-focused supervision to increase the training weight assigned to moments such as finding answer-relevant evidence or redirecting a stalled search. They apply this emphasis during supervised fine-tuning and reinforcement learning, teaching the model to recognize productive discoveries and ineffective search paths.
The ablations favor context control
The paper reports the following AREX-Base results. Scores use each benchmark's evaluation protocol and should be compared under matching tool access, search budgets, harnesses, and evaluator settings.
| Benchmark | Reported score | Paper's note |
|---|---|---|
| GAIA | 85.4 | Tool-using general assistant tasks |
| BrowseComp | 82.5 | Multi-hop web research |
| DeepSearchQA | 89.9 | Deep-search question answering |
| HLE with tools | 52.4 | Expert-level questions with tool access |
| WideSearch-en | 82.0 | Best reported score according to the authors |
On BrowseComp, the same AREX harness paired with Gemini Pro 3.1 reached 85.9, compared with 82.5 for AREX-Base. That result provides one view of the harness's contribution, although the underlying models and training recipes differ.
The BrowseComp ablations give context management a measurable role. Autonomous Context Updating added 11.8 percentage points over a standard linear-search baseline. Combining context updating with the outer self-improvement loop produced a cumulative gain of 22.9 points. The full figure reflects both mechanisms, while the isolated effect of the outer loop depends on the complete ablation configuration.
AREX-Turbo also outperformed models nearly ten times its size on most benchmarks in BAAI's evaluation. In a direct comparison, the fine-tuned 4B agent beat an untuned Qwen3.5-35B model on five of six benchmarks. The comparison measures the complete agent and training recipe, with parameter count representing only one variable.
Where recursive verification pays off
Constraint-heavy tasks provide the clearest use case because each verification pass can produce a specific research target. Suitable workloads include:
- Multi-source questions that require several facts to agree.
- Literature reviews with explicit inclusion criteria.
- Multi-hop browsing across documents, organizations, or time periods.
- Evidence synthesis that must preserve intermediate findings across long searches.
Single lookups and short reasoning chains receive less benefit because the outer loop, confidence checks, and context compression add tool calls and latency. Recursive verification becomes useful when unresolved checks yield concrete sub-questions that can guide another search.
Five implementation choices to test
Teams implementing a similar pattern can separate AREX's design into concrete engineering decisions:
- Represent requirements explicitly. Store each constraint separately so the verifier can evaluate it and generate a focused follow-up query.
- Carry forward compact evidence. Preserve supported claims, their sources, and unresolved gaps instead of replaying the full transcript.
- Expose compression as an agent action. Train the model to decide when compaction is needed and which information must remain available.
- Weight decisive trajectory steps. Emphasize evidence discoveries, search redirections, and stopping decisions during training.
- Calibrate cost and confidence. Set stopping thresholds on held-out tasks and measure quality gains against added search calls, latency, and token use.