Papers for

image synthesis developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

FuseReg improves image generation by fusing encoder layers flexibly

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Abstract: Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
When making computers create images, it's tricky to decide which parts of the visual understanding to use, because some parts help keep fine detail while others improve overall quality. The authors propose FuseReg, a method that trains the system to use different combinations of these parts randomly, so it learns to handle all of them well. This makes the image generator better at both recreating detailed images and producing high-quality new images without needing to change the original visual encoder.
Open → 2609.31620v1

Curvature-aware correction improves stability in diffusion image generation

Geometry-Aware Diffusion Guidance via Curvature-Adaptive Tubular Correction

Abstract: Gradient-guided diffusion samplers provide flexible priors for inverse problems and conditional generation, but strong guidance can move the sampling trajectory into regions where the learned score is poorly supported. Existing tangent-projection strategies limit first-order departure from an iso-density surface, yet discard potentially useful normal motion and overlook the second-order departure induced by tangent motion on a curved surface. We introduce curvature-adaptive tubular correction (CAT), a training-free plugin that regulates both effects within a shared, noise-dependent geometric budget. CAT decomposes the guidance gradient into normal and tangent components, charges normal displacement at first order and tangent displacement according to directional curvature, and obtains their jointly optimal magnitudes from a one-dimensional dual equation. Armijo backtracking calibrates the resulting finite step against the actual guidance objective, while matrix-free directional derivatives avoid constructing the full score Jacobian. We establish local guarantees for the tubular approximation, uniqueness of the correction, and sufficient objective decrease. Across seven inverse problems on FFHQ and ImageNet, CAT improves the evaluated pixel- and latent-space host samplers, with particularly consistent gains in perceptual metrics. It also improves black hole reconstruction on InverseBench and yields the lowest FID among the compared methods at every tested classifier-free guidance scale, while maintaining stable saturation and contrast. These results support curvature-aware tubular control as a reusable mechanism for stabilizing diffusion guidance.

Fri 18 SeptComputer Vision and Pattern Recognition
The gist
When computer programs create images by gradually refining noise, they follow clues called gradients that guide the process. However, if these clues are too strong, the program can get confused and produce worse images. The authors came up with a way to carefully adjust these clues by considering the shape of the data’s surface, making the guidance more stable and accurate. Their new method, called curvature-adaptive tubular correction, improved image quality in several test problems, including making sharper faces and clearer natural images.
Open → 2609.21251v1