Reinforcement learning cuts queries for black box vision attacks
Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision
Machine LearningCryptography and SecurityComputer Vision and Pattern Recognition
Summary
Neural networks used in computers to understand images can be tricked by tiny changes that humans hardly notice but cause wrong answers. Most ways to fool these networks need full access to their inner workings, which is unrealistic. The authors propose a method inspired by reinforcement learning that works without knowing the model details, by smartly trying small changes and learning from feedback. Their approach needs fewer tries to fool models and works well even on models trained to resist attacks.
What this means in practice
- •For security engineers: Conduct efficient black-box adversarial testing to evaluate and improve the robustness of computer vision systems against unknown threats.
- •For ai safety teams: Use reinforcement learning techniques to simulate realistic black-box attacks for validating defenses in safety-critical AI applications.
Authors
Florian Krone, Elena Hoemann, Sven Hallerbach
Abstract
Neural networks, both convolution or transformer based, are essential for modern computer vision systems. However, they are vulnerable to small perturbations, almost imperceptible to humans, which significantly alter the model's prediction. These adversarial attacks are often considered to be a significant threat to the implementation of neural networks in safety-critical applications. Most attacks utilize the white-box threat model and therefore require full access to the target model, making them unrealistic to use in practice. We propose a novel approach under the more realistic black-box threat model that utilizes concepts from reinforcement learning to optimize perturbations with a non-differentiable target model. Reinforcement learning algorithms have already been optimized to be query efficient, making them an ideal starting point when designing black-box adversarial attacks. We show the success of our reinforcement learning inspired black-box adversarial attack (RIBA) in generating adversarial perturbations using only a small number of queries to the target model, by comparing it to state of the art attacks on different models on the Cifar10 and ImageNet data sets. RIBA takes $25.4\%$ fewer median queries to generate attacked images against a ResNet-18 on Cifar10 and $22.5\%$ fewer median queries to fool a Vit-B/16 model on ImageNet. Additionally, we demonstrate that RIBA can match the performance of white-box attacks on an adversarially trained model.