Mlp adapters reduce cost of processing all visual tokens in language models
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Multimodal language models that understand images and text often spend a lot of computing power analyzing many visual pieces repeatedly. The authors found that instead of removing some visual parts, they can keep all pieces but process them more efficiently using simple math tricks called low-rank approximations and small neural networks called MLPs. Their method, called delta-Vision, saves computation while still keeping important visual information to answer questions or generate text accurately. This approach works well on images and videos, making models faster without losing important details.
What this means in practice
- •For ai system engineers: Reduce computational costs in multimodal AI systems that handle images and text by replacing repeated token processing with lightweight reconstruction modules.
- •For video content platform developers: Maintain full visual information when processing video frames for tasks like captioning or retrieval while improving inference efficiency.
Authors
Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang, Yiwen Guo, Soujanya Poria
Abstract
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $δ$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $δ$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.