Open-Source VisionHOPE Backbone Hits 85.6% Accuracy by Adapting How it Learns
A visual backbone that rewrites its own learning rule as it scans an image, hitting 85.6% ImageNet top-1 and 51.8 mIoU on ADE20K.
- VisionHOPE is a new visual backbone whose memory and learning rule co-evolve while scanning an image.
- Built on Nested Learning's self-referential construction with five coupled memories: content, key, value, learning rate, retention.
- Stability guaranteed via soft injection cap plus spectral clamp, proven non-expansive along each scan.
- Hits 85.6% ImageNet-1K top-1, 50.5 COCO Box AP, and 51.8 ADE20K mIoU with the Base model.
- MIT-licensed PyTorch code, Tiny/Small/Base weights on Hugging Face, drop-in SRNL and operator modules.
- Paper on arXiv; requires Linux, CUDA 12.1, and PyTorch 2.1.
VisionHOPE lets a visual backbone adapt how it learns
VisionHOPE is a newly released open-source visual backbone that changes its internal update controls while processing each image. Its authors published full PyTorch code, pretrained weights, and training recipes for ImageNet-1K, COCO, and ADE20K in a GitHub repository under the MIT license. The package includes Tiny, Small, and Base models alongside reusable operator modules.
The design represents its update process through five temporary memories covering content, keys, values, learning rate, and retention. These states evolve during a forward pass. Here, “learning” means updating temporary within-image memory; gradient-based checkpoint training remains a separate process.
Update rules become state
Modern visual architectures use several forms of input-conditioned computation. Convolutional Neural Networks aggregate local neighborhoods, Vision Transformers build input-dependent attention maps, State-Space Models vary state transitions, and Test-Time Training layers update an inner learner while processing an example.
VisionHOPE draws on Nested Learning, a framework that treats a model as coupled learning processes. Its temporary content memory changes alongside the variables that control how strongly new information enters and how much previous information remains.
A Transformer keeps its learned projection matrices fixed during inference even though its attention maps depend on the input. VisionHOPE extends input conditioning to optimizer-like controls inside the operator, allowing update size and retention to change token by token.
Five memories move together
During each scan, the operator maintains five coupled states with distinct roles:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.