GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

Computer Vision and Pattern Recognition

Summary

The authors developed GramLoop, a method to make a frozen DINOv3 model better at recognizing objects and scenes when the input images change in unexpected ways, like being noisy or altered. Instead of retraining the model, they add extra steps during prediction inside the model's existing layers, checking if the new computations stay consistent with the original model’s relationships between image patches. GramLoop selectively accepts improvements at each step to refine the model’s predictions without breaking its understanding of the image structure. Tests show that GramLoop boosts performance on several benchmarks with shifted image data while maintaining accuracy on normal data.

Authors

Yang Chen, Canyu Shen, Xinzhe Rao, Yuanyi Yan, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu

Abstract

We aim to improve frozen DINOv3 dense-prediction models under distribution shift by adding inference computation inside the visual backbone, without changing model weights, task adapters, or prediction heads. The challenge is that repeated transformer-block computation must refine dense features without disrupting the pairwise patch relations that DINOv3 uses to preserve spatial structure. We introduce GramLoop, a training-free framework that replays a short transformer window and controls each replay through final-layer cosine-Gram consistency. Each proposal is propagated through the frozen suffix, measured against the standard DINOv3 trajectory, and accepted through a patchwise gate at the replay-window endpoint. Across object detection and semantic segmentation under corruptions, perturbations, and natural shifts, GramLoop improves all five shifted benchmarks over the paired DINOv3 baseline. On COCO-O, it improves mAP by +0.252 and Effective Robustness by +0.250, while preserving clean ADE20K performance. Code will be released.