Zhipu AI's GLM-OCR Beats Gemini and GPT-5.2 at 260x Smaller

Zhipu's 0.9B parameter open-source OCR model tops OmniDocBench V1.5 at 94.62, beating Gemini 3 Pro and Qwen3-VL-235B on document parsing.

·
·
Zhipu AI's GLM-OCR Beats Gemini and GPT-5.2 at 260x SmallerPRO
Read2 min
TypeRepo
SubtopicOcr · Vision Language
  • GLM-OCR is a 0.9B open-weights OCR model from Zhipu AI, released under Apache 2.0 code and MIT weights on GitHub.
  • Scores 94.62 on OmniDocBench v1.5, beating Gemini 3 Pro (90.3), GPT-5.2 (85.4), and Qwen3-VL-235B (89.2).
  • Architecture pairs a CogViT visual encoder with a GLM-0.5B decoder plus PP-DocLayout-V3 for layout.
  • Trained with Multi-Token Prediction loss and reinforcement learning across all OCR tasks.
  • Throughput reaches 1.86 PDF pages/sec, with hosted API pricing near $0.03 per million tokens.
  • Ships an SDK with pip install glmocr, supporting vLLM, SGLang, Ollama, and MLX backends.

A document parsing model small enough to run on a laptop just outscored the largest frontier vision-language models on the benchmark everyone quotes. GLM-OCR, released by Zhipu AI under Apache 2.0 (with weights under MIT), packs state-of-the-art OCR into 0.9 billion parameters and ships with a one-line SDK, cloud API, and self-host paths through vLLM, SGLang, Ollama, and MLX.

Despite its 0.9B parameter count, GLM-OCR posts the highest Overall score on OmniDocBench v1.5 at 94.62, ahead of specialized competitors like PaddleOCR-VL-1.5 (94.50) and MinerU2.5 (90.67), and ahead of far larger general VLMs including Qwen3-VL-235B (89.15) and Gemini-3 Pro (90.33). A roughly 260x smaller model is beating the biggest closed systems on their own turf.

Small model, big table

Here is how the headline benchmarks stack up across the OCR-focused evaluations reported in the technical report:

BenchmarkGLM-OCR (0.9B)Gemini-3 ProGPT-5.2PaddleOCR-VL-1.5
OmniDocBench v1.594.690.385.494.5
OCRBench (Text)94.091.983.775.3
UniMERNet (formulas)96.596.490.596.1
TEDS_TEST (tables)86.081.867.683.3

The pattern holds across the report: GLM-OCR tops or ties the frontier on text, formulas, and tables at a fraction of the compute. It also sets a clear open-source lead in key information extraction, scoring 93.7 on Nanonets-KIE and 86.1 on Handwritten-KIE, narrowing the gap with closed systems like Gemini-3 Pro. The one place it does not sweep is PubTabNet, where MinerU 2.5 still leads.

How a 0.9B model gets there

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads