Baidu's TURA Beats a 671B AI Model Using a Tiny 4B Agent

Baidu's production AI search combines RAG with agentic tool-use via MCP servers, DAG planning and a distilled 4B executor serving millions.

·
·
Baidu's TURA Beats a 671B AI Model Using a Tiny 4B AgentPRO
Read2 min
TypePaper
SubtopicRag · Tool Use · Mcp
  • Baidu's TURA merges RAG with agentic tool-use, serving tens of millions of users in production.
  • Three stages: intent-aware MCP server retrieval, DAG-based task planner, and a distilled agent executor.
  • Distilled Qwen3-4B hits 88.3% tool-calling accuracy vs 82.4% for the 671B Deepseek-V3 teacher.
  • DAG parallel planning cuts latency 44% on complex multi-hop queries with no accuracy loss.
  • Live A/B test: Session Success Rate rose from 55.1% to 64.0% over the RAG baseline.
  • MCP-Bench released with 10,683 annotated queries across 61 MCP servers.

Traditional AI search chokes when you ask it something like "is there a bullet train from Beijing to Shanghai tomorrow with seats left?" The retrieval index only knows about static web pages, not the live inventory sitting behind a ticketing API. Baidu's search team has published TURA, a three-stage framework that fuses RAG with agentic tool calling and has been running in production since May 2025, serving tens of millions of users.

Side-by-side comparison of TURA calling a ticket API versus RAG returning stale web results

The static web problem

Current RAG systems, optimized for retrieving from pre-indexed web documents, are fundamentally incapable of accessing dynamic, real-time information that is not present on a static webpage but must be generated via interaction. Flight seats, train schedules, live inventory, weather at a specific hour, none of that lives in a crawled corpus. Search engines historically patched this with hand-curated components like Google's OneBox or Baidu's Aladdin platform, but those approaches rely on brittle, hand-crafted integrations that are difficult to scale.

TURA replaces those bespoke boxes with a unified agentic pipeline where every information source, static web index or dynamic API, is wrapped as a Model Context Protocol (MCP) server. The core objective is stated formally: maximize answer quality subject to a strict latency budget, since this is a live search product where seconds matter.

Three stages, one pipeline

TURA framework overview showing intent-aware retrieval, DAG planner, and distilled executor

The architecture has three modules that hand off to each other:

  1. Intent-Aware MCP Server Retrieval. An LLM decomposes the raw query into atomic sub-intents, then dense-retrieves candidate MCP servers for each.
  2. DAG-based Task Planner. A planner LLM builds a directed acyclic graph of sub-tasks so independent calls run in parallel.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads