How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks

2026-08-17Sound

Sound
AI summary

The authors study how well different neural audio watermarking methods hold up against attempts to remove them. Instead of blindly testing many attacks, they analyze how the watermarks are embedded by comparing pairs of clean and watermarked audio. This lets them apply targeted removal attacks that are more effective and efficient. They find that some watermarks hidden in certain parts of the audio are easy to remove without much quality loss, while those embedded in deeper features are much harder to break. Their technique can also identify which watermarking method was used with good accuracy.

neural audio watermarkingembedding domainaudio watermark removalpayloadlatent domainmagnitude domaincarrier domainPESQaccuracy-quality trade-offstructural probes
Authors
Likhith Kumara
Abstract
Neural audio watermarks are increasingly used to attribute and detect AI-generated speech, so their practical value rests on how cheaply an adversary can remove them. Robustness is usually measured by running a fixed battery of distortions blindly against every scheme. We instead make removal diagnostic: from a few clean/watermarked pairs we compute cheap structural probes that reveal where a watermark sits in the signal (its embedding domain), then apply a single domain-matched attack rather than a blind sweep. We further summarize each scheme with one threshold-free fragility score, the area under its accuracy-versus-quality trade-off, which an accuracy-only benchmark cannot provide. Across ten watermarking schemes the probes separate fragile from robust marks: for magnitude and carrier-domain watermarks a single matched attack erases the payload (WavMark, SilentCipher, audiowmark) or removes the detection flag (AudioSeal) at high objective quality (PESQ >= 3.6), whereas latent-domain marks (VoiceMark, WMCodec, AlignMark, AWARE) resist every training-free attack we apply. The same pair-only probe signatures also identify which watermarking scheme is present (84% over ten schemes).