Baidu's Qianfan-OCR Beats 235B Models at Document Parsing With Just 4B
Baidu's 4B vision-language model replaces multi-stage OCR pipelines with a single pass, topping OmniDocBench at 93.12 while running on one GPU.
PRO- Baidu released Qianfan-OCR, a 4B end-to-end document intelligence VLM under Apache 2.0.
- Scores 93.12 on OmniDocBench v1.5 and 880 on OCRBench, leading all end-to-end models.
- Layout-as-Thought uses think tokens to emit bounding boxes and reading order before generating Markdown.
- Beats Gemini-3.1-Pro and Qwen3-VL-235B on KIE with 87.9 average across five benchmarks.
- Runs at 1.024 pages/sec on a single A100 with W8A8 quantization, deployable via vLLM.
- Supports 192 languages, tables, formulas, charts, handwriting, and prompt-driven field extraction.
Document OCR has long been a Rube Goldberg machine: one model finds text boxes, another recognizes characters, a third handles tables, and an LLM stitches it all together. Baidu's Qianfan team just released a model that collapses that entire chain into a single 4B-parameter vision-language model, and the numbers suggest the pipeline era may finally be ending.
Qianfan-OCR is a 4B-parameter end-to-end document intelligence model developed by the Baidu Qianfan Team that unifies document parsing, layout analysis, and document understanding within a single vision-language architecture. Unlike traditional multi-stage OCR pipelines that chain separate modules for layout detection and text recognition, Qianfan-OCR performs direct image-to-Markdown conversion and supports prompt-driven tasks like table extraction and document question answering.
Why pipelines were breaking
The problem with traditional OCR stacks is not accuracy on plain text, it's what gets thrown away between stages. Traditional OCR pipelines rely on separate modules for layout detection and text recognition, often resulting in spatial reasoning failures where visual context like chart axis relationships is discarded. By contrast, Qianfan-OCR's unified architecture maintains this context, allowing it to succeed where two-stage systems scored 0.0 on CharXiv benchmarks.
That zero is not a typo. When you extract text from a bar chart and hand it to an LLM as a string, you lose the axis positions, the color mapping, and which label goes with which bar. The chart becomes unreadable. An end-to-end VLM keeps the pixels and the reasoning in the same model, so spatial relationships survive.

Layout-as-Thought: giving the model a scratchpad
The trade-off with going end-to-end is that you lose explicit layout output. Pipeline systems naturally produce bounding boxes and reading order because those are intermediate artifacts. A pure image-to-Markdown model just gives you the text. Qianfan-OCR's core trick is what the authors call Layout-as-Thought.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.