Open-Source VisionHOPE Backbone Hits 85.6% Accuracy by Adapting How it Learns

A visual backbone that rewrites its own learning rule as it scans an image, hitting 85.6% ImageNet top-1 and 51.8 mIoU on ADE20K.

·
·
·
Open-Source VisionHOPE Backbone Hits 85.6% Accuracy by Adapting How it LearnsPRO
  • VisionHOPE is a new visual backbone whose memory and learning rule co-evolve while scanning an image.
  • Built on Nested Learning's self-referential construction with five coupled memories: content, key, value, learning rate, retention.
  • Stability guaranteed via soft injection cap plus spectral clamp, proven non-expansive along each scan.
  • Hits 85.6% ImageNet-1K top-1, 50.5 COCO Box AP, and 51.8 ADE20K mIoU with the Base model.
  • MIT-licensed PyTorch code, Tiny/Small/Base weights on Hugging Face, drop-in SRNL and operator modules.
  • Paper on arXiv; requires Linux, CUDA 12.1, and PyTorch 2.1.

VisionHOPE lets a visual backbone adapt how it learns

VisionHOPE is a newly released open-source visual backbone that changes its internal update controls while processing each image. Its authors published full PyTorch code, pretrained weights, and training recipes for ImageNet-1K, COCO, and ADE20K in a GitHub repository under the MIT license. The package includes Tiny, Small, and Base models alongside reusable operator modules.

The design represents its update process through five temporary memories covering content, keys, values, learning rate, and retention. These states evolve during a forward pass. Here, “learning” means updating temporary within-image memory; gradient-based checkpoint training remains a separate process.

Update rules become state

Modern visual architectures use several forms of input-conditioned computation. Convolutional Neural Networks aggregate local neighborhoods, Vision Transformers build input-dependent attention maps, State-Space Models vary state transitions, and Test-Time Training layers update an inner learner while processing an example.

VisionHOPE draws on Nested Learning, a framework that treats a model as coupled learning processes. Its temporary content memory changes alongside the variables that control how strongly new information enters and how much previous information remains.

A Transformer keeps its learned projection matrices fixed during inference even though its attention maps depend on the input. VisionHOPE extends input conditioning to optimizer-like controls inside the operator, allowing update size and retention to change token by token.

Progression from convolutional networks to nested learning with increasing within-image adaptation
VisionHOPE extends input-conditioned computation to the controls governing temporary memory updates.

Five memories move together

During each scan, the operator maintains five coupled states with distinct roles:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads