Allen AI's Molmo2-ER Beats GPT-5 and Gemini on Robot Vision at 4B Parameters

Allen AI's 4B vision language model beats GPT-5 and Gemini Robot-ER 1.5 Thinking on 9 of 13 embodied reasoning benchmarks.

·
·
Allen AI's Molmo2-ER Beats GPT-5 and Gemini on Robot Vision at 4B ParametersPRO
  • Allen AI released Molmo2-ER, a 4B open-weight VLM specialized for embodied and spatial reasoning.
  • Scores 63.8% average across 13 benchmarks, beating GPT-5 (57.9) and Gemini Robot-ER 1.5 Thinking (61.3).
  • Built on Qwen3-4B plus SigLIP2, fine-tuned from Molmo2 with a +17 point gain.
  • Trained via specialize-then-rehearse on a 3.3M-sample corpus (SAT, RoboPoint, RefSpatial, VSI-590K, others).
  • Serves as the vision backbone for Ai2's MolmoAct2 robot action model.
  • Apache 2.0 licensed, runs in Transformers, vLLM, and SGLang; code at github.com/allenai/molmo2.

Allen AI has quietly shipped a small vision language model that punches well above its weight for robotics work. Molmo2-ER is a 4B parameter model tuned for the perception skills robot policies actually need: pointing at exact pixels, matching objects across camera views, and reasoning about geometry over video. It is Apache 2.0 licensed, available on Hugging Face, and works out of the box with Transformers, vLLM, and SGLang.

A 4B model that outscores the giants

Molmo2-ER outperforms every open-weight baseline along with the strongest closed-source models, including Gemini Robotics-ER 1.5 Thinking and GPT-5, on 9 of 13 established embodied reasoning benchmarks. It posts an overall average of 63.8%, a 17-point jump over the Molmo2 starting point.

On the same 13-benchmark suite (Point-Bench, RefSpatial, BLINK, CV-Bench, ERQA, EmbSpatial, MindCube, SAT, VSI-Bench, and others), the comparison looks like this:

  • Molmo2-ER: 63.8
  • Gemini Robotics-ER 1.5 Thinking: 61.3
  • Qwen3-VL-8B: 61.0
  • GPT-5: 57.9
  • Gemini 2.5 Pro: 57.1

Why a specialized backbone was needed

Existing VLM backbones are optimized for semantic image understanding rather than the metric, geometric, and temporally grounded reasoning required for robot control. A general VLM can tell you there is a mug on a table, but it struggles to say exactly which pixel to grasp, how far away the object sits, or which item in camera B corresponds to the one in camera A.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads