When a multimodal model gets a visual question wrong, is it because it failed to see something, or because it failed to reason about it? That distinction sounds simple, but almost no existing benchmark actually separates the two. PerceptionBench, just released by Moonshot AI, is built specifically to answer that question, and the results are sobering.

The Benchmark Nobody Built (Until Now)

Most visual benchmarks test end-to-end performance: you give a model an image and a question, and you score the answer. The problem is that a wrong answer could mean the model didn't see the object, didn't understand the spatial relationship, hallucinated something that wasn't there, or simply failed to reason through the problem. These are very different failure modes, and conflating them makes it nearly impossible to know what to fix.

Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. PerceptionBench takes a different approach entirely.

PerceptionBench task categories compared to existing benchmarks, showing 10 atomic perception categories with sample questions

Built From Failures, Not Assumptions

Rather than deciding upfront what "visual perception" means, the Moonshot AI team did something more rigorous: they looked at where frontier models actually break. By attributing frontier-model failures across 42 existing benchmarks to their earliest visual cause, they distill 10 perceptual capabilities and 3,000 verified questions, each answerable by looking, with no reasoning or outside knowledge required. This bottom-up, failure-driven taxonomy is the core methodological contribution.

Alpha Signal

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves