SafeStyle improves style transfer while avoiding content leaks

SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization

Computer Vision and Pattern Recognition

Summary

When changing a picture’s style using another image and text instructions, it’s hard to keep the new style without accidentally copying parts of the original picture. The authors developed SafeStyle, a method that carefully adds style details without letting unwanted original content leak through. It works without extra training, by separating and controlling style and content information to find a better balance. Tests show that SafeStyle keeps style strong while reducing accidental content copying.

What this means in practice

  • For graphic designers: Create stylized images from text prompts guided by style references without unwanted content copying from the style image.
  • For video game artists: Use controlled style editing tools to apply complex textures or shapes while maintaining character or scene identity.

Authors

Zhangping Yang, Min Li, Song Yan, Rong Gao, Xinliang Bi, Guanye Xiong, Yujie He

Abstract

Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8\% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.