Linear vision transformers close gap with pretrained softmax weights
Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
Computer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
Summary
Training new types of vision models called linear ViTs is usually hard and less effective compared to traditional ones that use softmax attention. The authors show that simply copying the attention parts from pretrained softmax models doesn't work well because those parts behave differently. Instead, copying the other parts (called MLP weights) and teaching the attention in a new way helps linear ViTs perform just as well or better. This method works for different model sizes and datasets, making it easier to reuse existing models when switching to linear attention.
What this means in practice
- •For computer vision engineers: Improve training efficiency of linear vision transformer models by reusing pretrained weights from softmax transformers with targeted weight copying and distillation.
- •For machine learning platform teams: Integrate better initialization methods for linear attention models to speed up deployment and optimization of large-scale vision transformer workflows.
Authors
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, Wu Liu, Xi Peng, Chun Jian Ho, Hongyuan Zhu
Abstract
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.