Papers for

video game artists

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Style transfer method reduces content leakage while keeping style intact

CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

Abstract: Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Style transfer tries to make a picture look like it's painted in the style of another image, but often bits of the original picture sneak through, making the final image look messy. The authors found that pushing too hard to remove these bits often ruins the style, and trying to keep the style strong brings back unwanted content. They created a new method called CLeaR that balances this by carefully separating content and style using multiple image analysis tools without extra training. This approach results in better style matching and less leakage in the generated images.
Open → 2609.38136v1

SafeStyle improves style transfer while avoiding content leaks

SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization

Abstract: Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8\% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
When changing a picture’s style using another image and text instructions, it’s hard to keep the new style without accidentally copying parts of the original picture. The authors developed SafeStyle, a method that carefully adds style details without letting unwanted original content leak through. It works without extra training, by separating and controlling style and content information to find a better balance. Tests show that SafeStyle keeps style strong while reducing accidental content copying.
Open → 2609.21242v1

Vehicle shape and size recovered from limited images using priors

Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views

Abstract: High-fidelity vehicle assets are essential for controllable traffic scene generation, particularly for synthesizing rare and safety-critical long-tail scenarios. However, reconstructing a reusable vehicle representation from in-the-wild onboard images remains challenging for two reasons. First, image-to-3D generation methods generally produce models without reliable metric scale. Second, onboard cameras usually observe only one side of a target vehicle, making conventional multi-view reconstruction incomplete on unobserved regions. To solve these problems, we propose a feed-forward vehicle asset reconstruction method, which leverages two complementary priors to reconstruct 3D Gaussian representations for vehicles using sparse one-sided observations. To achieve metric-scale reconstruction, a visual foundation model is first utilized to serve as a visual prior for Gaussian initialization. The Gaussian attributes are then estimated by a learnable encoder-decoder module. A symmetry-aware cloning strategy is presented to complete the unobserved side directly in Gaussian space, which exploits the bilateral structure of vehicles as a geometric prior. Experiments on the public dataset demonstrate that the proposed method significantly outperforms existing approaches in both vehicle asset completeness and geometric accuracy.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
It is difficult to create accurate 3D models of vehicles from just a few photos taken from one side because the scale is unclear and the other side is missing. The authors use visual clues from a foundation model to start the shape and size, then apply a special method that copies the seen side to the unseen side, since most vehicles are symmetrical. This approach makes it possible to build a complete and correctly sized 3D vehicle model even from just one viewpoint. Their tests show this method works better than others in creating full and precise vehicle shapes.
Open → 2609.08841v1