Research

Mininglamp Releases Open-Source VisionHOPE Backbone

Mininglamp Technology and its partners have open-sourced VisionHOPE, a new vision backbone that uses adaptive memory to outperform traditional vision transformers on high-resolution tasks.

AlphaSignal1 day agoResearch
Image: AlphaSignal

Mininglamp Technology, alongside the Chinese Academy of Sciences' Institute of Automation and other collaborators, has released VisionHOPE, an open-source visual backbone. Published under an MIT license, the model introduces an adaptive runtime memory that evolves dynamically as it processes an image. Unlike conventional vision transformers, state-space models, or convolutional networks that keep their learned parameters static during a forward pass, VisionHOPE adjusts its internal states token by token.

The architecture is built on a framework called Nested Learning, which treats the model as a series of interconnected learning processes. VisionHOPE implements this by coupling five distinct memories: content, key, value, learning rate, and retention. During inference, the model uses a delta-rule-style operation to update these memories based on the difference between predicted and target values. To ensure these self-referential updates remain stable and non-expansive, the system employs a soft injection cap and a spectral clamp.

This design achieves linear compute and memory scaling relative to token count. On Nvidia A100 GPUs, VisionHOPE outperformed established models like DeiT and Vim when handling high-resolution images. In benchmark testing, the model achieved Top-1 accuracy scores on ImageNet-1K of 84.1 percent for the Tiny variant, 85.2 percent for the Small variant, and 85.6 percent for the Base variant. The researchers also reported strong performance on the COCO dataset for object detection and instance segmentation, as well as the ADE20K dataset for semantic segmentation.

For computer vision practitioners, the release provides a highly flexible and efficient alternative to standard vision transformers. Because the model parameters remain fixed during inference while only the per-input runtime states evolve, practitioners get the benefits of an active optimization process without permanently altering the base model. The open-source release includes code and weights for classification, detection, and segmentation tasks, allowing developers to immediately integrate VisionHOPE into existing pipelines.

This is our own summary of reporting by AlphaSignal

More in Research