Image experts improve video diffusion models with preserved motion

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Making videos using AI models is often slow and needs lots of video data. The authors found a way to use image-based AI models, which are cheaper and easier to run, to improve video-making AI without losing the natural motion in videos. They created a tool called MILD that connects image and video models, letting the video model learn from the image model but still keep smooth movement. This method works better than previous techniques that only used video models as teachers.

What this means in practice

  • For video content creators: Improve AI-generated video quality by incorporating detailed image-based features while maintaining natural video motion and consistency.
  • For mobile app developers: Use cheaper image-based AI models to enhance video generation quality in apps without the heavy computational cost of large video models.

Authors

Bingqing Jiang, Li Luo, Zichao Yu, Yujin Han, Zhaolong Su, Difan Zou

Abstract

On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.