Baidu's MapAgent Cuts Human Map Editing Across 360 Cities to 5%
Baidu's MapAgent uses a Judge-Planner-Worker agent loop to fix lane-map errors at scale, pushing production automation past 95% across 360+ cities
- MapAgent wraps a lane-map prediction backbone in a Judge-Planner-Worker agent loop to catch and fix specification violations automatically.
- A VLM-based Judge, trained with SFT then GRPO reinforcement learning, inspects BEV images and draft lane vectors together to diagnose rule violations.
- The system only activates on low-confidence map tiles, keeping city-scale throughput practical.
- Deployed in Baidu Maps across 360+ cities, pushing production automation above 95%.
- Gains are largest on hard long-tail cases: worn markings, occlusions, unusual intersections -- exactly where human editors spend most time.
- No public code or model weights released; paper accepted to KDD '26.
MapAgent, a new agentic framework from Baidu, shows that you can take an existing lane-map prediction model and wrap it in a structured agent loop to dramatically cut the human editing burden -- without slowing down city-scale production. The system is already live inside Baidu Maps, handling lane-level map generation for over 360 cities across China.
The map-making bottleneck nobody talks about
Lane-level maps are core infrastructure for autonomous driving, advanced driver assistance, and lane-level navigation, providing centimeter-level priors on road geometry, lane topology, and traffic control. The problem is keeping them fresh and correct at scale.
The dominant paradigm relies on specialized survey vehicles equipped with high-precision LiDAR and other sensors, followed by extensive manual annotation -- a traditional approach plagued by prohibitive operational costs, labor-intensive processes, and slow update cycles.
Modern deep learning has made a dent. Methods such as HDMapNet, VectorMapNet, MapTR, and MapTRv2 convert multi-sensor inputs into bird's-eye-view (BEV) features and directly decode vectorized polylines or topology, replacing much of the manual mapping pipeline. But these models have a blind spot: they learn mapping rules implicitly from training data, so they can't reliably enforce the explicit specifications that production maps must satisfy.
End-to-end vectorized mapping methods typically treat mapping specifications and traffic regulations as implicit, dataset-dependent supervision. In complex scenes -- worn or missing markings, occlusions -- correct lane configurations are often under-determined by visual evidence alone, making specification violations a major source of human post-editing.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.