Mininglamp's Open-Source VisionHOPE Beats Vision Transformers With Adaptive Memory
A new vision backbone rewrites its own learning rules while scanning an image, hitting strong ImageNet, COCO, and ADE20K numbers with linear-time cost.
- VisionHOPE is a new open-source vision backbone whose memory and learning rule co-evolve during a single forward pass.
- Built on Nested Learning, it couples five memories: content, key, value, learning rate, retention.
- A soft injection cap plus spectral clamp provably keep the self-referential updates non-expansive.
- Hits 84.1/85.2/85.6 Top-1 on ImageNet-1K at Tiny/Small/Base, plus strong COCO and ADE20K results.
- Linear compute and memory in token count, outperforming DeiT and Vim at high resolution on A100.
- MIT-licensed weights and code cover classification, detection, and segmentation.
VisionHOPE is an open-source visual backbone whose runtime memory adapts as it scans each image. Most ConvNets, vision transformers, and state-space models keep their learned parameters fixed throughout a forward pass. VisionHOPE adds five coupled memories that update token by token, giving the model an input-specific learning process while preserving linear scaling with token count.
Researchers from Mininglamp Technology, the Chinese Academy of Sciences’ Institute of Automation, and collaborating institutions describe the architecture in a new paper. The MIT-licensed release includes Tiny, Small, and Base checkpoints for ImageNet-1K classification, COCO object detection and instance segmentation, and ADE20K semantic segmentation.
Five memories inside one pass
VisionHOPE builds on Nested Learning, a framework that treats a model as interconnected learning processes that compress incoming context into internal states. The vision architecture implements that idea with five memories whose outputs influence one another’s next update.
| Memory role | Function during a scan |
|---|---|
| Content | Stores information accumulated from visual tokens. |
| Key | Generates representations used to address the content memory. |
| Value | Generates the information written into memory. |
| Learning rate | Controls the size of each token-level update. |
| Retention | Controls how much existing state survives each update. |
Each memory is represented by small matrices or vectors and updated with a delta-rule-style operation. At each token, the model compares the value predicted from the current memory with the generated target value, then uses that difference to adjust the state. Outputs from the learning-rate and retention memories determine how strongly the other memories change and how much prior information they keep.
Checkpoint parameters remain fixed during inference. The evolving quantities are per-input runtime states, initialized for each directional scan and discarded after the forward pass. The result resembles a small optimization process embedded inside inference without permanently changing the downloaded model.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.