Video super-resolution improves low-quality videos using local frame context

LoCoVSR: Local Context Diffusion Posterior Sampling for Video Super-Resolution

Computer Vision and Pattern Recognition

Summary

Video super-resolution means turning low-quality videos into sharper, clearer high-quality ones, but it’s a tricky problem with many possible solutions. The authors created LoCoVSR, a new method that looks at small groups of nearby low-resolution video frames to better guess the details in the high-resolution video. Their method avoids complicated steps like estimating motion between frames and works well on long videos without losing quality over time. They tested LoCoVSR on face video data and showed it makes videos look sharper and smoother over time compared to similar recent methods.

What this means in practice

  • For video streaming platforms: Upscale low-resolution video streams in real-time without motion estimation to enhance user viewing quality efficiently.$Commercial implications: Enables streaming services to offer better video quality on low-bandwidth connections with a cost-effective, scalable solution.
  • For video editing professionals: Restore old or noisy videos by enhancing fine details and smoothness without needing complex motion tracking between frames.

Tested on one dataset.

Authors

Matan Ben Chorin, Michael Elad

Abstract

Video super-resolution (VSR) is an ill-posed inverse problem that aims to reconstruct a high-resolution (HR) video from a noisy, low-resolution (LR) version of it. We present LoCoVSR, a diffusion-based VSR framework that leverages pixel-space denoising diffusion probabilistic models. LoCoVSR integrates the Diffusion Posterior Sampling technique with spatio-temporal context learning, operating in a moving-average form. A localized window of adjacent LR frames is used for recovering each center frame, while applying a shared noise trajectory across all frames. The localized windowing enables processing of long videos without length limitations, supports parallel inference, and prevents error accumulation that may occur in recursive processing. Unlike prior methods, LoCoVSR offers a simple yet very effective VSR solution, avoiding explicit optical flow estimation, or information loss caused by latent space processing. Trained on the VFHQ face dataset, LoCoVSR achieves accurate, temporally consistent and high-quality upscaling with competitive results against recent diffusion-based VSR approaches.