Video generation models gain reusable tools for faster diverse tasks

LongLive-Plug: Once-for-All Distillation for Video Generation

Computer Vision and Pattern Recognition

Summary

Video generation models often need extra training when adapted for different jobs, which takes time. The authors created LongLive-Plug, a way to train helpful add-ons just once on a main model. These add-ons can then be plugged into many different video models without retraining, making tasks like faster sampling and longer video correction easier. This approach was tested on many models and tasks, showing the add-ons work well across different settings.

What this means in practice

  • For video production teams: Use reusable model adapters to speed up video generation and improve long video quality without retraining for each new task.
  • For robotics developers: Deploy trained plug-in modules on robotic video generation models to achieve efficient sampling and long-context correction across varied scenarios.

Authors

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

Abstract

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.