Information limits of watermarking images from pretrained AI generators

On the Information-Theoretic Limits of Latent-Space Watermarking Through Pretrained Generators

Information TheoryMachine Learning

Summary

This paper studies how much secret information, or watermarks, can be hidden in images created by AI generators when using a secret key and a message. The authors look at the limits on reliably hiding watermarks while keeping the images realistic and unchanged in style. They find mathematical bounds for how much information can be encoded and how this changes when an attacker tries to remove the watermark by regenerating the image. Their work helps understand how watermarks can be hidden in AI-created images and how robust these marks are to attacks.

What this means in practice

  • For digital content platforms: Determine fundamental limits on embedding hidden watermarks in AI-generated images to protect content ownership under realistic usage and attack scenarios.
  • For multimedia security teams: Assess how repeated image regeneration techniques reduce watermark detectability and plan watermark designs that maintain persistence against such attacks.

A theory result. No direct application yet.

Authors

Jinwan Jeon, Minju Lee, Sung Hoon Lim

Abstract

We study latent-space watermarking through a pretrained generator using a prescribed latent-to-output stochastic mapping, called the renderer. A watermark encoder selects the latent input using a message and secret key. For every message and semantic context, the released output must have exactly the desired conditional output distribution. For finite alphabets, we derive rate--key inner and outer bounds and characterize the coding and coordination requirements for realizing watermark communication through the prescribed latent interface. When the target output distribution of the generator uniquely determines the corresponding latent input distribution through the renderer, a strengthened converse yields the capacity region; the same region governs explicit preservation of the pretrained latent distribution. We extend the analysis to general jointly Gaussian models and identify a sufficient statistic of the latent that captures both the watermark-bearing information available at the generated output and the latent coordination required to preserve its target distribution. For the vector Gaussian model, we further characterize the optimal allocation of the secret-key resource across the resulting modes. Finally, we turn to an emerging robustness threat that is particularly natural in generative watermarking: an adversary can regenerate the released sample to obtain a fresh realization of the same underlying content while attenuating or destroying the embedded watermark. We incorporate this robustness axis into our framework and characterize the one-pass compound capacity of the scalar Gaussian model when the semantic context is known to the encoder but hidden from the detector, while the regeneration attack may depend on that context. Extending the analysis to multiple rounds of repeated canonical regeneration, we characterize the resulting watermark-capacity decay.