Vision language models vulnerable to hidden backdoor attacks in image frequencies

FreqDoor: A Hidden Trojan in the Frequency Domain for Backdoor Attacks on Vision-Language Models

Computer Vision and Pattern Recognition

Summary

Multimodal AI systems that generate text from images can be secretly tricked using hidden signals. The researchers found a method called FreqDoor that hides these secret triggers in parts of the image you can't see, called the frequency domain. These triggers don’t change the visible image or text but still manipulate what the AI says in response. The method worked very well on popular models, making them respond incorrectly almost all the time, while still producing text that seems normal. This reveals a new kind of security risk in AI models that combine images and language.

Vision-language modelsBackdoor attackFrequency domainAmplitude spectrumPhase informationImage captioningVisual question answeringBLIP-2InstructBLIPLLaVA

Authors

Yasir Arafat Prodhan, Sadad Hasan, Mohammed Imamul Hassan Bhuiyan

Abstract

Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. However, their multimodal nature makes them persistently vulnerable to backdoor attacks. Existing backdoor triggers for VLMs are either spatial, textual, or bimodal, which may yield localized or recognizable trigger patterns. In this work, we explore a different attack surface and propose \ textsc {FreqDoor}, a training-time backdoor attack that implants triggers in the frequency domain. \ textsc {FreqDoor} mixes amplitude-spectrum components from a trigger-source image selectively while preserving the phase of a clean image to generate a spatially distributed and visually imperceptible trigger without modifying the textual input. We evaluate the attack on BLIP-2, InstructBLIP, and LLaVA for image captioning and visual question answering. On Flickr8k, \ textsc {FreqDoor} achieves attack success rates of $99.6\%$, $99.8\%$, and $98.4\%$ on the three models, respectively, while preserving the semantic quality of the generated captions. On VQAv2, the corresponding attack success rates are $99.6\%$, $92.4\%$, and $79.6\%$.