Layer aware position embeddings improve visual token pruning in multimodal models

Layer-Aware Position Embeddings for Visual Token Pruning in Multimodal Large Language Models

Computer Vision and Pattern Recognition

Summary

Multimodal large language models process images by breaking them into many small pieces called visual tokens, which takes a lot of computing power. To speed things up, some tokens can be removed, but this can confuse the model because the positions of the remaining tokens get reassigned in different ways, each with drawbacks. The authors found that some parts of the model care more about keeping tokens in their original positions to understand images well. They designed a method that changes how positions are reassigned depending on the layer, improving overall image understanding while making the model faster.

What this means in practice

  • For multimodal ai engineers: Reduce computing costs in large multimodal models while maintaining image understanding by applying layer-aware position embedding during token pruning.
  • For mobile app developers: Deploy faster and more efficient AI-powered image captioning or question answering by integrating this token pruning strategy to reduce resource use.

Authors

Yahong Wang, Zhangkai Ni, Juncheng Wu, Yuyin Zhou, Ying Wen, Lianghua He

Abstract

Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained tokens using either sparse or continuous position embeddings, each introducing distinct limitations. Sparse position embeddings tend to decrease the attention value allocated to visual tokens, thereby degrading the perception capability of MLLMs, whereas continuous position embeddings disrupt the original spatial correspondence of visual tokens, leading to weakened grounding capability. To mitigate this issue, we perform layer-wise analysis of the language decoder and observe that intermediate layers play a critical role for maintaining the grounding capability of MLLMs under token pruning. Based on this observation, we propose a layer-aware position embedding strategy, which switches to sparse position embeddings at grounding-sensitive layers while maintaining continuous position embeddings elsewhere. Extensive experiments across representative pruning methods and diverse benchmarks demonstrate that our approach improves the comprehensive multimodal performance of pruned MLLMs compared with standard sparse and continuous position embeddings.