Structured frequency attacks reveal robust model weaknesses

Frame the adversary: a structure-aware attack methodology

Machine Learning

Summary

Some tricks can fool AI by making tiny changes to pictures that humans barely notice. This paper shows a new way to create these tricky changes by thinking about how images behave in terms of their hidden frequency patterns, rather than just pixels. The authors use math to carefully shape these changes, making attacks that are strong and work across different AI designs. Their method helps us understand AI weaknesses better, beyond quick fixes tied to one specific system.

What this means in practice

  • For security teams: Create stronger tests for AI systems by using frequency-based attacks that reveal deeper model weaknesses across different architectures.
  • For image recognition engineers: Improve AI robustness by understanding and defending against structured frequency perturbations that generalize beyond specific model designs.

Authors

Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre

Abstract

Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.