Baidu's Unlimited-OCR Reads Entire Documents Without Chunking or Stitching

Baidu's Unlimited-OCR uses a novel attention mechanism to parse entire multi-page documents in a single pass, eliminating the chunking bottleneck that plagues current OCR pipelines.

·
·
Baidu's Unlimited-OCR Reads Entire Documents Without Chunking or StitchingPRO
Read2 min
TypeRepo
SubtopicOcr
  • One-shot multi-page parsing: Baidu's Unlimited-OCR processes entire PDFs in a single forward pass, no chunking required.
  • New attention mechanism: Reference Sliding Window Attention (R-SWA) keeps KV cache size constant regardless of document length, fixing the core scaling problem.
  • Strong benchmark results: Achieves 93.92% on OmniDocBench v1.6 SOTA, +6% over the DeepSeek-OCR baseline it builds on.
  • Efficient model size: 3B total parameters, only 500M active per pass (MoE architecture), runs on a single GPU.
  • MIT licensed and available now: Weights on Hugging Face, supports Transformers and SGLang inference backends.
  • Generalizable architecture: R-SWA is applicable beyond OCR to ASR, translation, and any reference-based long-horizon generation task.

Every serious OCR pipeline today has a dirty secret: it's just a for-loop. You slice a document into individual pages, run each through a model, then stitch the outputs back together , hoping nothing breaks at the seams. Tables that span pages get split. Context from page 3 is invisible on page 10. It's an engineering workaround that the field has quietly accepted as the norm. Unlimited-OCR, a new open-source model from Baidu, is built to make that workaround obsolete.

The problem with how OCR works today

Modern end-to-end OCR models like DeepSeek-OCR use a large language model as the decoder, which gives them strong language priors and impressive accuracy on individual pages. But as the output sequence lengthens, the accumulated KV cache drives up memory consumption and progressively slows down generation. The KV cache here is the memory the model keeps of everything it has already generated , and with standard attention, it grows unboundedly as the document gets longer. The practical result: you can't just feed a 40-page PDF to these models and get a coherent output. You have to chunk it.

Most open-source OCR pipelines force you to cut input into individual pages, run each through a model separately, then stitch the outputs back together. That stitching is where things go wrong , cross-page tables break, reading order gets confused, and the model has no memory of what it already transcribed.

A human analogy that actually holds up

This stands in stark contrast to humans, who exhibit no such decline in efficiency during long-horizon copying tasks. Unlimited OCR is a model designed to emulate human parsing working memory. When a person copies a book, they don't memorize every word they've written. They keep track of where they are in the source, remember the last few words they wrote, and let older content fade from working memory. That's exactly the intuition behind the new architecture.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads