Spectral features detect attacks on image super-resolution models

Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features

Computer Vision and Pattern Recognition

Summary

Image super-resolution models can be secretly manipulated by hackers to trick later processing steps. The authors show a way to spot these attacks by looking at patterns in the frequencies of the model's internal weights. They use a special kind of analysis called radially-averaged power spectral density and train a machine learning detector to recognize when a model has been tampered with. Their approach works better than other similar methods in most tests, especially by focusing on high-frequency details in the model.

What this means in practice

Authors

Emma J. Reid, Haley Duba-Sullivan, Tony G. Allen

Abstract

The integration of deep learning models into image preprocessing pipelines such as super-resolution introduces a largely unexplored attack vector for adversaries targeting downstream tasks. To ensure trustworthiness of critical imaging pipelines, we must be able to detect adversarial behavior within preprocessing models. In this paper, we propose a spectral-based detection method for identifying adversarial attacks embedded in super-resolution model weights. More specifically, we use the radially-averaged power spectral density as a discriminative feature to train an extreme gradient boosting (XGBoost) detector, demonstrating detectability of model-level threats in super-resolution networks. We further benchmark our detector against magnitude- and phase-based Fourier spectrum detectors, evaluating each method across a range of training and cross-architecture scenarios. Our proposed detector out-performs the comparison detectors in most of these scenarios and indicates that high-frequency features are most informative for detecting AdvSR attacks across SR architectures.