Baidu Rebuilds Search With Four AI Agents, Beats RAG by 13%
Baidu Search unveils a four-agent blueprint that plans, executes, and reflects on multi-step queries, moving beyond one-shot retrieve-then-generate RAG.
PRO- Baidu Search proposes a four-agent search paradigm: Master, Planner, Executor, and Writer, replacing linear RAG pipelines. Paper
- The Planner decomposes complex queries into a DAG of sub-tasks bound to tools via the Model Context Protocol.
- Dynamic Capability Boundary narrows thousands of MCP tools down to about a dozen candidates per query.
- DRAFT refines tool docs iteratively; COLT retrieves complete tool sets using bipartite query-scene-tool graphs.
- Live A/B test on Baidu Search shows +1.85% DAU, +1.04% PV, and -1.45% change query rate versus legacy.
- Human eval reports +13% Normalized Win Rate on complex queries and +5% on moderately complex ones.
Baidu Search has published a detailed blueprint for what it calls the AI Search Paradigm, a multi-agent system designed to replace the linear retrieve-then-generate pipeline that dominates today's RAG-based search assistants. Instead of a single model doing everything, four specialized LLM-powered agents coordinate to break queries apart, call tools, verify results, and write the final answer, and the whole thing is already running on live traffic at Baidu.
The paper is less about a single benchmark win and more about a full architectural rethink, complete with algorithms for planning, tool retrieval, reinforcement learning, adversarial robustness, and inference-time efficiency. The authors argue that current RAG systems function as passive retrievers rather than active problem solvers, and their fix is a cognitive-style workflow of agents that reason before they retrieve.
Why one-shot RAG keeps breaking
The motivating example is a question that looks easy but isn't: Who was older, Emperor Wu of Han or Julius Caesar, and by how many years? No single document explicitly compares them. A vanilla RAG pipeline retrieves once, sees the two names, and hallucinates. Even advanced approaches like ReAct or RQ-RAG fail because they rely on in-context reasoning without invoking a calculator, so they may guess who's older but can't compute the exact age gap.

The paradigm treats this as a planning problem. Complex queries need to be decomposed, dependencies tracked, tools bound to sub-tasks, results verified, and only then synthesized into a coherent answer.
The four agents and three team modes
The system is built around four roles, each backed by an LLM:
- Master: reads the query, judges its complexity, assembles the right team, and watches for failures. If something goes wrong, it triggers reflection and re-planning.
- Planner: only invoked for complex queries. It decomposes the query into a Directed Acyclic Graph (DAG) of sub-tasks and binds each node to a tool.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.