Z.ai's Tiny GLM-OCR Beats Models 260x Larger at Reading Documents
Z.ai released a 0.9B-parameter OCR model that tops OmniDocBench V1.5 with a 94.62 score, beating models 260x larger.
PRO- Z.ai released GLM-OCR, a 0.9B-parameter multimodal document OCR model, MIT-licensed.
- Scores 94.62 on OmniDocBench V1.5, ranking #1 and beating Qwen3-VL-235B and Gemini-3 Pro.
- Uses CogViT encoder plus GLM-0.5B decoder, trained with Multi-Token Prediction and full-task RL.
- Two-stage pipeline with PP-DocLayoutV3 parses tables, formulas, handwriting, and seals into Markdown or JSON.
- Runs on vLLM, SGLang, Ollama, and Transformers; hits 1.86 PDF pages per second.
- Over 3 million Hugging Face downloads in the first month after release.
Z.ai has quietly upended the document understanding leaderboard with GLM-OCR, a compact multimodal model that turns messy PDFs, invoices, formulas, and handwritten notes into structured Markdown, JSON, and LaTeX. It does this with only 0.9 billion parameters, small enough to run on a single consumer GPU, and it ships fully open source under the MIT license.
In February 2026, the model topped the OmniDocBench V1.5 leaderboard with a score of 94.62, beating Qwen3-VL-235B (260 times larger) and closing in on closed-source systems like Gemini-3 Pro at 90.33. Released by Z.ai (Zhipu AI), it racked up over 3 million downloads in its first month on Hugging Face.

What's actually inside the 0.9B
GLM-OCR is a multimodal OCR model built on the GLM-V encoder-decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The stack combines the CogViT visual encoder pre-trained on large-scale image-text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder.
Two design choices matter here. Multi-Token Prediction trains the decoder to predict several future tokens at once instead of one at a time, which speeds up both training and inference. Rather than dumping every page straight into the VLM, GLM-OCR uses a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, so tables, formulas, and body text can be routed and recognized in parallel.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.