SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

2026-07-20Sound

Sound
AI summary

The authors address the problem of tracking synthetic speech by creating a new watermarking method called SSTMark. Unlike traditional methods that hide watermarks in the actual audio signals, SSTMark hides the watermark in the words spoken, making it easier to detect even after the audio is altered. Their tests show SSTMark is better at detecting watermarks, especially when the speech undergoes common edits or compression. This means SSTMark could help in reliably identifying synthetic speech.

speech watermarkingsynthetic speechsignal-level representationssemantic-level watermarkingtext watermarkingAudioMarkBenchfalse positive ratespeech compressionspeech processingdetection rate
Authors
Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien Lu
Abstract
As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow. Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1\%, SSTMark improves the average detection rate by 4.6\% and 16.9\% on signal-processing edits and compression edits, respectively.