Defend vision language models from attacks without retraining

Backdoor as Probe: Test-Time Adversarial Defense for CLIP

Computer Vision and Pattern Recognition

Summary

Vision-language models like CLIP can be tricked by tiny changes in images that cause mistakes. Instead of trying to ignore these changes, the authors use them as signals to detect attacks by adding a hidden backdoor trigger inside the model. When an attack happens, this backdoor gets strongly activated, helping the system spot and fix the problem at test time without needing to retrain. Their method greatly improves the model's accuracy against attacks while keeping normal performance and runs faster than other defense methods.

What this means in practice

  • For security teams in ai software: Detect and fix adversarial attacks on vision-language models at test time without retraining, improving security and reliability.
  • For machine learning engineers: Speed up deployment of vision-language models by adding defenses that require no expensive retraining and run faster than existing methods.

Authors

Zhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen, Shiguang Shan, Xilin Chen

Abstract

Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose \emph{Backdoor as Probe} (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models. Project page: https://robin-wzq.github.io/Backdoor-as-Probe/