Visual token compression method cuts compute while keeping accuracy

Beyond Selection: Token Parameterization for Extreme Visual Token Compression

Computer Vision and Pattern RecognitionMachine Learning

Summary

Models that understand images and language often use many small visual pieces called tokens. When there are too many tokens, it becomes slow and expensive to process them. The authors studied how to compress these tokens into fewer, more efficient ones without hurting the model’s ability to understand images well. They designed a new method called Braco that smartly transforms and compresses tokens to keep important details and speed up processing. Tests show Braco can greatly reduce computation while keeping accuracy high compared to older compression methods.

What this means in practice

  • For vision-language engineers: Speed up vision-language model training and inference by reducing visual token processing without losing accuracy.
  • For mobile app developers: Deploy image understanding features efficiently on mobile devices by cutting down computation with compressed visual tokens.

Authors

Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo

Abstract

Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under $23\times$--$64\times$ compression and remains competitive at $144\times$, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.