Feature rearrangement improves single image generation quality and structure

FRPSS: Feature Rearrangement in Pre-Shape Space for Single-Image Generation

Computer Vision and Pattern Recognition

Summary

Generating new images from just one original picture is tricky because the new images can lose important shapes or look too similar. The paper presents a new method that rearranges features from the original image in a special way to keep the overall structure intact while allowing variety in details. This method reduces mix-ups in the image layout during generation. The authors also created tools to help stylize images based on this approach, showing better performance than previous methods.

What this means in practice

  • For graphic designers: Create diverse variations of a single image that maintain overall structure for creative projects and stylization.
  • For digital artists: Produce high-quality single-image based artwork with better spatial coherence and varied details for digital art creation.

Authors

Yuexing Han, Haoxuan Zhang, Bing Wang

Abstract

Generative models trained on a single image often struggle to balance global structural integrity and local diversity. Existing single-image generation methods commonly rely on random noise to drive the generation process and lack explicit global structural constraints, making the generated results prone to spatial structural misalignment when structural variations occur. To address the issue, Feature Rearrangement in Pre-Shape Space for Single-Image Generation (FRPSS) is proposed in this paper. The core of FRPSS is the Manifold Structural Rearrangement with Feature Augmentation on Geodesic Surface (MSR-FAGS) module. MSR-FAGS replaces the randomly initialized features of the low-scale generator with rearranged Pre-Shape features and uses the features to guide image generation at subsequent scales, thereby reducing the risk of structural misalignment. To support downstream tasks such as stylization, a Scale-adaptive Sliding-window Patch Extraction (SSPE) strategy is further designed, and a directional Contrastive Language-Image Pre-training supervision module with SSPE (CLIP-SSPE) is constructed. Qualitative and quantitative experiments demonstrate that FRPSS achieves the best Single Image Fréchet Inception Distance (SIFID) scores on all three datasets while maintaining competitive Learned Perceptual Image Patch Similarity (LPIPS). Further qualitative experiments verify the effectiveness of FRPSS across multiple downstream tasks with the CLIP-SSPE module.