Zhipu AI's GLM-OCR Beats Gemini and GPT-5.2 at 260x Smaller
Zhipu's 0.9B parameter open-source OCR model tops OmniDocBench V1.5 at 94.62, beating Gemini 3 Pro and Qwen3-VL-235B on document parsing.
PRO- GLM-OCR is a 0.9B open-weights OCR model from Zhipu AI, released under Apache 2.0 code and MIT weights on GitHub.
- Scores 94.62 on OmniDocBench v1.5, beating Gemini 3 Pro (90.3), GPT-5.2 (85.4), and Qwen3-VL-235B (89.2).
- Architecture pairs a CogViT visual encoder with a GLM-0.5B decoder plus PP-DocLayout-V3 for layout.
- Trained with Multi-Token Prediction loss and reinforcement learning across all OCR tasks.
- Throughput reaches 1.86 PDF pages/sec, with hosted API pricing near $0.03 per million tokens.
- Ships an SDK with
pip install glmocr, supporting vLLM, SGLang, Ollama, and MLX backends.
A document parsing model small enough to run on a laptop just outscored the largest frontier vision-language models on the benchmark everyone quotes. GLM-OCR, released by Zhipu AI under Apache 2.0 (with weights under MIT), packs state-of-the-art OCR into 0.9 billion parameters and ships with a one-line SDK, cloud API, and self-host paths through vLLM, SGLang, Ollama, and MLX.
Despite its 0.9B parameter count, GLM-OCR posts the highest Overall score on OmniDocBench v1.5 at 94.62, ahead of specialized competitors like PaddleOCR-VL-1.5 (94.50) and MinerU2.5 (90.67), and ahead of far larger general VLMs including Qwen3-VL-235B (89.15) and Gemini-3 Pro (90.33). A roughly 260x smaller model is beating the biggest closed systems on their own turf.
Small model, big table
Here is how the headline benchmarks stack up across the OCR-focused evaluations reported in the technical report:
| Benchmark | GLM-OCR (0.9B) | Gemini-3 Pro | GPT-5.2 | PaddleOCR-VL-1.5 |
|---|---|---|---|---|
| OmniDocBench v1.5 | 94.6 | 90.3 | 85.4 | 94.5 |
| OCRBench (Text) | 94.0 | 91.9 | 83.7 | 75.3 |
| UniMERNet (formulas) | 96.5 | 96.4 | 90.5 | 96.1 |
| TEDS_TEST (tables) | 86.0 | 81.8 | 67.6 | 83.3 |
The pattern holds across the report: GLM-OCR tops or ties the frontier on text, formulas, and tables at a fraction of the compute. It also sets a clear open-source lead in key information extraction, scoring 93.7 on Nanonets-KIE and 86.1 on Handwritten-KIE, narrowing the gap with closed systems like Gemini-3 Pro. The one place it does not sweep is PubTabNet, where MinerU 2.5 still leads.
How a 0.9B model gets there
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.