Moonshot AI's PerceptionBench Reveals No Frontier Model Sees Past 60%
Moonshot AI's PerceptionBench reveals no frontier model clears 60% accuracy on pure visual perception tasks, exposing a critical gap in how we evaluate multimodal AI.

- PerceptionBench released: Moonshot AI launches a 3,000-question benchmark isolating pure visual perception from reasoning in multimodal models.
- No model clears 60%: Across 16 frontier models including GPT-5.6-Sol, Kimi K3, and Claude-Fable-5, none exceed 59.7% accuracy on atomic perception tasks.
- Failure-driven taxonomy: 10 perception categories were derived by tracing model failures across 42 existing benchmarks to their earliest visual cause.
- Hallucination is the weakest link: Perception-related hallucination is the lowest-scoring category on average; GPT-5.6-Sol scores only 26.9% despite leading overall.
- Models often guess, not perceive: A significant share of correct answers don't hold up on repeated questioning, revealing fragile pattern-matching rather than genuine perception.
- Free and open: Dataset on Hugging Face (CC-BY-NC-4.0), eval code on GitHub (Apache 2.0), works with any OpenAI-compatible endpoint.
When a multimodal model gets a visual question wrong, is it because it failed to see something, or because it failed to reason about it? That distinction sounds simple, but almost no existing benchmark actually separates the two. PerceptionBench, released by Moonshot AI, is built specifically to answer that question, and the results are uncomfortable.
The Gap in Visual Benchmarking
Most visual benchmarks test end-to-end performance: give a model an image and a question, then score the answer. A wrong answer could mean the model didn't see the object, misread a spatial relationship, hallucinated something absent, or failed to reason through the problem. These are distinct failure modes, and conflating them makes it nearly impossible to know what to fix.
Holistic evaluations conflate perceptual errors with reasoning failures, while application-driven benchmarks cover only narrow, fragmented domains. PerceptionBench takes a different approach entirely.
Built From Actual Model Failures
Rather than deciding upfront what "visual perception" means, the Moonshot AI team looked at where frontier models actually break. By tracing failures across 42 existing benchmarks to their earliest visual cause, they distilled 10 perceptual capabilities and 3,000 verified questions, each answerable by looking, with no reasoning or outside knowledge required. That bottom-up, failure-driven taxonomy is the core methodological contribution.
The 10 atomic categories are: