Fine-tuning probabilistic image translators with reward adjustments

Tilted Schrödinger Bridge Matching

Machine Learning

Summary

Sometimes computers need to change pictures from one style to another without paired examples, which can be tricky. The authors worked on improving a mathematical method called Schrödinger bridges that helps with this kind of unpaired image transformation. They introduced a way to update these models after training to better match certain goals or preferences, like making digits or faces look a certain way. Their approach carefully adjusts outputs while keeping important properties intact. They tested their method on datasets of digits and face images.

What this means in practice

  • For computer vision teams: Adjust image-to-image translation models after training to prioritize specific image features or user-defined rewards.
  • For machine learning engineers: Incorporate reward-based fine-tuning into probabilistic generative models to meet application-specific constraints or preferences.

Authors

Sergei Kholkin, Evgeny Burnaev, Alexander Korotin

Abstract

Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion models but underexplored for Schrödinger bridges. We introduce Tilted Schrödinger Bridge Matching (TSBM), a post-training method for fine-tuning a learned bridge $P$ between source $p_0$ and target $p_1$ toward a reward-tilted target $p_1^r\propto p_1e^r$, while preserving source $p_0$. We formulate this adaptation as alternating optimization initialized from $P$, provide theoretical justification, and derive a practical algorithm based on Adjoint Matching. We evaluate TSBM on unpaired image-to-image translation targeting digit properties in MNIST and facial attributes in CelebA.