PhoenixSR improves real-world image super-resolution efficiency and quality
PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution
Computer Vision and Pattern Recognition
Summary
Enhancing low-quality images to look realistic and detailed is hard, especially in real-world situations where images can be very complex. The authors introduce PhoenixSR, a new way to teach faster, simpler image-enhancing models by learning from a slower but smarter model type called diffusion models. PhoenixSR transfers knowledge without needing the slow parts when actually improving images, keeping the process fast and efficient. Experiments show that PhoenixSR makes image results look better while still keeping details accurate.
What this means in practice
- •For mobile app developers: Create smartphone apps that enhance photo resolution quickly while maintaining realistic details using PhoenixSR-trained models.$Commercial implications: Enables efficient, high-quality image enhancement features in consumer apps that run on devices without heavy computational resources.
- •For video streaming platforms: Deploy efficient super-resolution models to improve video quality on low-resolution streams without increasing processing lag.
Authors
Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha
Abstract
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing diffusion-free SR networks without introducing diffusion components at inference time? To this end, we propose PhoenixSR, a generative heterogeneous distillation framework that transfers diffusion priors to independently designed feed-forward SR networks through score-based distribution matching. Rather than aligning heterogeneous features or imitating sampled diffusion outputs, PhoenixSR uses the pretrained diffusion model as distribution-level supervision, while paired SR supervision preserves reconstruction fidelity. To make distribution matching effective for fidelity-sensitive SR, we introduce Heterogeneous Distribution Adaptation, which adapts the target score to the SR domain, improves tracking of the evolving student distribution, and anchors training with paired supervision. We further employ Directional Reliability Weighting, a lightweight residual-consistency-based reweighting strategy that reduces unstable distributional guidance. All diffusion-related components are removed after training, leaving the original student architecture and inference cost unchanged. Experiments on three SR benchmarks and six feed-forward backbones, including SwinIR, HAT, Real-ESRGAN, and SeeMoRe, show consistent perceptual improvements with largely preserved reconstruction fidelity.