Likelihood score approximation improves inverse problem restoration quality

Livin' on a Prior: Likelihood Score Approximation for Inverse Problems

Machine Learning

Summary

Solving problems where you want to recover original data from altered or degraded versions is tricky, especially when the exact changes are unknown. The authors present a method called Likelihood Score Approximation (LSA) that uses a fixed, pretrained model and learns from a small amount of paired examples to better guess what the original data looked like. This approach works well even with very limited training data and can improve restoration of both images and speech. It also allows changing the base model after training without losing performance.

What this means in practice

Authors

Rostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann

Abstract

Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting that spans both regimes: unknown degradations can be learned from few paired examples, while known degradations can be learned from self-generated samples. We introduce Likelihood Score Approximation (LSA), a generative framework that keeps a pretrained unconditional model fixed and learns an observation-conditioned model that approximates the likelihood score from paired samples. Within a conditional stochastic-interpolant framework, LSA can be trained in either score or velocity coordinates, independently of the unconditional model's native parameterization, and supports both deterministic and stochastic sampling. We further show empirically that the prior model can be swapped post-training while keeping the same LSA model. Across speech and image inverse problems, LSA operates effectively even at roughly 0.01% of the full training dataset. On the ImageNet-256 benchmark it achieves competitive or better restoration quality than strong posterior-sampling baselines while requiring up to several orders of magnitude fewer network evaluations.