Watermarking masked diffusion language models using token pairs for detection
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Machine LearningComputation and LanguageCryptography and Security
Summary
Text generated by some AI writing models is difficult to watermark because they fill in words in any order, not left to right. The authors found that simpler watermarking methods make certain words appear more often and can be tricked by attackers. They developed TANGO, which connects each new word to a nearby already-known word to hide a secret signal in word pairs, making the watermark harder to detect by outsiders. Tests on real models show TANGO can spot watermarked text reliably while resisting attacks that fooled older methods.
What this means in practice
- •For language model providers: Add secret watermarks to text from masked diffusion models that survive edits and avoid easy detection.$Commercial implications: Enables service providers to prove ownership of AI-generated text and prevent forgery in products using these models.
- •For content moderation teams: Detect AI-generated text watermarked by masked diffusion models to improve authenticity verification of online content.
Authors
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen
Abstract
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.