Vision language models vulnerable to hidden backdoor attacks in image frequencies
FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models
Computer Vision and Pattern Recognition
Summary
Multimodal AI systems that generate text from images can be secretly tricked using hidden signals. The researchers found a method called FreqDoor that hides these secret triggers in parts of the image you can't see, called the frequency domain. These triggers don’t change the visible image or text but still manipulate what the AI says in response. The method worked very well on popular models, making them respond incorrectly almost all the time, while still producing text that seems normal. This reveals a new kind of security risk in AI models that combine images and language.
Vision-language modelsBackdoor attackFrequency domainAmplitude spectrumPhase informationImage captioningVisual question answeringBLIP-2InstructBLIPLLaVA
Authors
Yasir Arafat Prodhan, Sadad Hasan, Mohammed Imamul Hassan Bhuiyan
Abstract
Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. However, their multimodal nature makes them persistently vulnerable to backdoor attacks. Existing backdoor triggers for VLMs are either spatial, textual, or bimodal, which may yield localized or recognizable trigger patterns. In this work, we explore a different attack surface and propose \ textsc {FreqDoor}, a training-time backdoor attack that implants triggers in the frequency domain. \ textsc {FreqDoor} mixes amplitude-spectrum components from a trigger-source image selectively while preserving the phase of a clean image to generate a spatially distributed and visually imperceptible trigger without modifying the textual input. We evaluate the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering. On Flickr8k, \ textsc {FreqDoor} achieves attack success rates of $99.6\%$, $99.8\%$, and $98.4\%$ on the three models, respectively, while preserving the semantic quality of the generated captions. On VQAv2, the corresponding attack success rates are $99.6\%$, $92.4\%$, and $79.6\%$.