Image editing models improve multi-step editing by learning from themselves

On-Policy Self-Distillation for Multi-Turn Image Editing

Computer Vision and Pattern RecognitionMachine Learning

Summary

Editing images using step-by-step instructions works well for one step but struggles when applied repeatedly to fix or change the image multiple times. The authors found that this happens because models train on perfect starting images but have to work with their own imperfect edits when used multiple times. They created a new training method that lets the model learn from its own outputs while still guided by a clean version, improving performance when editing images across many steps. To test this, they also built a benchmark with sequences of ten edits to better measure how models handle long editing tasks.

What this means in practice

  • For mobile app developers: Create photo editing apps that maintain quality even when multiple instructions are applied sequentially by users.$Commercial implications: This paper enables more reliable and user-friendly multi-step photo editing features for commercial mobile photography apps.
  • For digital content teams: Support iterative image creation workflows in design tools by improving models that handle many consecutive edits.

Authors

Liangbing Zhao, Le Zhuo, Mohamed Elhoseiny

Abstract

Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure to a train-test mismatch in the conditioning distribution: models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference time. To address this, we propose MT-OPSD, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations. We further introduce LME-Bench, a benchmark of 100 ten-turn editing sessions for evaluating long-horizon robustness. Experiments across three editing backbones show that MT-OPSD substantially improves long-horizon editing success and reduces multi-turn collapse while largely preserving single-turn editing quality.