FuseReg improves image generation by fusing encoder layers flexibly

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Computer Vision and Pattern Recognition

Summary

When making computers create images, it's tricky to decide which parts of the visual understanding to use, because some parts help keep fine detail while others improve overall quality. The authors propose FuseReg, a method that trains the system to use different combinations of these parts randomly, so it learns to handle all of them well. This makes the image generator better at both recreating detailed images and producing high-quality new images without needing to change the original visual encoder.

What this means in practice

  • For image synthesis developers: Improve image generators by training decoders to flexibly combine multiple encoder layers, yielding better reconstruction and generation quality without changing the encoder.
  • For computer vision engineers: Enhance robustness of visual feature fusion when building systems that reconstruct or generate images from pretrained encoders.

Authors

Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang

Abstract

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.