Compression introduces hidden risks in large vision-language models
Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models
Computer Vision and Pattern RecognitionArtificial IntelligenceCryptography and Security
Summary
Large vision-language models use compression to speed up image processing, but this can cause specific errors that don't appear when the model uses all its data. The authors identify cases where compression causes the model to fail even though it works fine without compression, which they call compression-specific failures. They designed a new attack, called CIRA, that finds images making the model fail only after compression, without needing access to all internal parts. The work also proposes a defense to reduce these failures, highlighting the importance of testing models both with and without compression to understand risks.
What this means in practice
- •For machine learning engineers: Identify and mitigate errors caused specifically by visual token compression in vision-language models to improve reliability in resource-constrained deployments.
- •For security teams: Use the CIRA attack method to test the vulnerability of compressed vision-language models without needing full model access, improving robustness evaluations.
Authors
Qiankun Li, Yuechen Zhang, Bowen Chen, Shilinlu Yan, Zhenhong Zhou, Kun Wang, Li Sun
Abstract
Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial input that remains correct under full-token inference but fails after compression, casting compression-induced risk as a paired failure attribution problem. Within a controlled diagnostic cohort, counterfactuals show that retained-set allocation causally changes compressed correctness and reveal a negative association between recovery and representation drift in displaced evidence. Motivated by these findings, we propose CIRA, a Compression-Induced Risk Attack for Large Vision-Language Models. Under a vision-encoder white-box setting, CIRA optimizes image perturbations through encoder-side objectives that manipulate token priorities across candidate compression budgets while preserving displaced evidence. CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across 12 dataset-compressor settings evaluated at four budgets, CIRA achieves a mean CSFR of 20.35% while limiting full-token attack success to 6.92%, with similar behavior on additional LVLM families. A cross-view selection-stabilization defense substantially suppresses CIRA, although Adaptive CIRA partially restores its effectiveness. These results show that compression-specific failures persist under restricted access and support paired evaluation of full-token and compressed inference for attributing risk to visual-token compression.