Multimodal language model efficiency improves with token pruning and skipping
SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Processing images and text together in large AI models can be very slow and costly because they handle a lot of unnecessary information and computations. The paper finds that looking at which visual tokens are important not just at the end but also in the middle layers of the model can help remove redundant data more effectively. It also shows that skipping some parts of the model during processing can save computation without losing much accuracy. The authors introduced SPIDER, a method that uses these ideas to make multimodal models faster while keeping most of their performance.
What this means in practice
- •For ai system engineers: Optimize multimodal AI deployments by reducing computational load while maintaining accuracy using SPIDER’s token pruning and layer skipping.
- •For mobile app developers: Implement more efficient image-and-text AI features on devices by applying SPIDER to reduce processing demand without major loss in quality.
Authors
Tianxiang Chen, Zhentao Tan, Zi Ye, Yue Wu, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Tao Gong, Qi Chu, Nenghai Yu, Xipeng Qiu, Jieping Ye
Abstract
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by $79\%$ while maintaining 96$\%$ of the baseline performance.