Style transfer method reduces content leakage while keeping style intact

CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

Computer Vision and Pattern Recognition

Summary

Style transfer tries to make a picture look like it's painted in the style of another image, but often bits of the original picture sneak through, making the final image look messy. The authors found that pushing too hard to remove these bits often ruins the style, and trying to keep the style strong brings back unwanted content. They created a new method called CLeaR that balances this by carefully separating content and style using multiple image analysis tools without extra training. This approach results in better style matching and less leakage in the generated images.

What this means in practice

  • For graphic designers: Generate stylized images that maintain the desired artistic style without copying unwanted details from the style source image.
  • For video game artists: Create game textures and assets that reflect specific art styles while avoiding the accidental reuse of objects or shapes from reference images.

Authors

Teng Zhou, Yunhao Chen

Abstract

Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.