Papers for

video streaming platforms

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Video super-resolution improves low-quality videos using local frame context

LoCoVSR: Local Context Diffusion Posterior Sampling for Video Super-Resolution

Abstract: Video super-resolution (VSR) is an ill-posed inverse problem that aims to reconstruct a high-resolution (HR) video from a noisy, low-resolution (LR) version of it. We present LoCoVSR, a diffusion-based VSR framework that leverages pixel-space denoising diffusion probabilistic models. LoCoVSR integrates the Diffusion Posterior Sampling technique with spatio-temporal context learning, operating in a moving-average form. A localized window of adjacent LR frames is used for recovering each center frame, while applying a shared noise trajectory across all frames. The localized windowing enables processing of long videos without length limitations, supports parallel inference, and prevents error accumulation that may occur in recursive processing. Unlike prior methods, LoCoVSR offers a simple yet very effective VSR solution, avoiding explicit optical flow estimation, or information loss caused by latent space processing. Trained on the VFHQ face dataset, LoCoVSR achieves accurate, temporally consistent and high-quality upscaling with competitive results against recent diffusion-based VSR approaches.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Video super-resolution means turning low-quality videos into sharper, clearer high-quality ones, but it’s a tricky problem with many possible solutions. The authors created LoCoVSR, a new method that looks at small groups of nearby low-resolution video frames to better guess the details in the high-resolution video. Their method avoids complicated steps like estimating motion between frames and works well on long videos without losing quality over time. They tested LoCoVSR on face video data and showed it makes videos look sharper and smoother over time compared to similar recent methods.
Open → 2609.32742v1

PhoenixSR improves real-world image super-resolution efficiency and quality

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing diffusion-free SR networks without introducing diffusion components at inference time? To this end, we propose PhoenixSR, a generative heterogeneous distillation framework that transfers diffusion priors to independently designed feed-forward SR networks through score-based distribution matching. Rather than aligning heterogeneous features or imitating sampled diffusion outputs, PhoenixSR uses the pretrained diffusion model as distribution-level supervision, while paired SR supervision preserves reconstruction fidelity. To make distribution matching effective for fidelity-sensitive SR, we introduce Heterogeneous Distribution Adaptation, which adapts the target score to the SR domain, improves tracking of the evolving student distribution, and anchors training with paired supervision. We further employ Directional Reliability Weighting, a lightweight residual-consistency-based reweighting strategy that reduces unstable distributional guidance. All diffusion-related components are removed after training, leaving the original student architecture and inference cost unchanged. Experiments on three SR benchmarks and six feed-forward backbones, including SwinIR, HAT, Real-ESRGAN, and SeeMoRe, show consistent perceptual improvements with largely preserved reconstruction fidelity.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Enhancing low-quality images to look realistic and detailed is hard, especially in real-world situations where images can be very complex. The authors introduce PhoenixSR, a new way to teach faster, simpler image-enhancing models by learning from a slower but smarter model type called diffusion models. PhoenixSR transfers knowledge without needing the slow parts when actually improving images, keeping the process fast and efficient. Experiments show that PhoenixSR makes image results look better while still keeping details accurate.
Open → 2609.30988v1