Improving autoregressive video generation by learning directly from reference videos

From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Generating videos one frame at a time often requires complex tools that can be slow and hard to train. The authors developed a way to learn video patterns directly from example videos, skipping some usual extra steps and models. Their method compares videos in a special way that measures differences without needing additional teacher models. This approach improves video quality while keeping the process efficient and allows learning new visual styles and concepts more easily.

What this means in practice

  • For video ai engineers: Build faster and higher-quality autoregressive video generators by training directly on example videos without auxiliary score models.
  • For visual effects studios: Create diverse visual styles and spatial effects in generated videos without relying on specialized teacher models for each style.

Authors

Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu

Abstract

Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.