Oxford and MIT Researchers Argue Vision Alone Could Unlock General AI
A group of vision researchers argues that pure vision, not language-tethered multimodal models, could be its own route to general intelligence.
PRO- Twenty-plus vision researchers propose Visual General Intelligence as a vision-first path to AGI.
- Authors include Andrew Davison, Jiajun Wu, Deva Ramanan, David Fouhey, Zhuang Liu, Yilun Du and Robert Geirhos.
- Paper questions whether pure vision research is advancing or being absorbed by language-centric VLMs.
- Draws GPT analogy: what emerges from scaling images, video, and geometry as primary modalities?
- Covers principles, input modalities, benchmarks, learning paradigms, and vision-language relationships rather than one definition.
- Ties into the CVPR 2026 VGI Workshop; no code or model released.
More than twenty computer vision researchers from Oxford, MIT, Stanford, CMU, Google DeepMind, Meta and AIST have released a white paper making an unusual argument for the current moment in AI: that vision on its own, without leaning on language as a crutch, may be a viable path toward general intelligence. They call this direction Visual General Intelligence, or VGI.
The paper, coordinated by Hirokatsu Kataoka and co-authored with names including Andrew Davison, Jiajun Wu, Deva Ramanan, David Fouhey, Zhuang Liu, Yilun Du, Robert Geirhos, Aditi Raghunathan, Christian Rupprecht and Yuki Asano, is positioned as a discussion document rather than a manifesto. Its aim is to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and how vision should relate to other modalities such as language when treated as the core.
Why vision researchers are getting nervous
The backdrop is uncomfortable for anyone whose career sits in pure computer vision. Large language models keep improving, and through Vision Language Models, and especially VLLMs where LLMs are fused with visual encoders, benchmark numbers keep climbing. That leaves an awkward question hanging over the field: is pure vision research still advancing on its own terms, or is it riding on progress made elsewhere?
Most headline progress on visual tasks now comes from bolting an image encoder onto a language model. The vision side of that pipeline often feels like a feature extractor feeding a much bigger linguistic brain. The white paper pushes back on that framing and asks what capabilities could emerge if vision were treated as the central substrate rather than a sensor for an LLM.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.