Baidu's Qianfan-OCR Beats Gemini and Qwen3-VL-235B With Just 4B Parameters
Baidu's new 4B parameter open-weight vision-language model tops OmniDocBench v1.5 and OCRBench while unifying parsing, layout, and understanding in one pass.
PRO- Baidu released Qianfan-OCR, a 4B end-to-end document intelligence VLM under Apache 2.0.
- Tops OmniDocBench v1.5 (93.12), OlmOCR Bench (79.8), and OCRBench (880) among end-to-end models.
- KIE average of 87.9 beats Gemini-3.1-Pro and Qwen3-VL-235B on structured extraction.
- Layout-as-Thought thinking mode emits bounding boxes and reading order before final text output.
- Runs at 1.024 pages per second on a single A100 with W8A8 quantization via vLLM.
- Supports 192 languages, chart QA, LaTeX formulas, complex tables, and handwriting in one pass.
Baidu's Qianfan team has released a document intelligence model that collapses the traditional OCR stack into a single 4B-parameter vision-language model. Qianfan-OCR is live on Hugging Face under Apache 2.0, and currently tops several end-to-end OCR leaderboards, beating models orders of magnitude larger.
The 4B checkpoint unifies document parsing, layout analysis, and document understanding inside one vision-language architecture. Where traditional multi-stage OCR pipelines chain separate layout detection, text recognition, and language comprehension modules, Qianfan-OCR performs direct image-to-Markdown conversion and supports prompt-driven tasks including structured document parsing, table extraction, chart understanding, document question answering, and key information extraction, all from the same weights.
The pipeline problem it kills
Classic document AI stacks look like Rube Goldberg machines. A detector finds text regions, a recognizer transcribes them, a layout analyzer figures out reading order, then an LLM tries to reason over the extracted text. Each stage compounds errors, and any spatial information that was not explicitly transcribed is gone forever.
The clearest evidence sits in chart understanding. Two-stage pipelines discard visual context like chart axis relationships during recognition, then score a flat 0.0 on CharXiv because the geometry of the chart is thrown away before reasoning begins. Qianfan-OCR's unified architecture keeps that context in the residual stream all the way through.
What is under the hood
The architecture borrows the multimodal bridging design from Qianfan-VL and glues together three parts:
- Vision encoder (Qianfan-ViT): An Any Resolution design that tiles images into 448 x 448 patches. Supports variable-resolution inputs up to 4K, producing up to 4,096 visual tokens per image to preserve small fonts and dense text.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.