Autoregressive video models generate longer videos with improved context retention

Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation

Computer Vision and Pattern Recognition

Summary

Many video generation methods struggle to keep details consistent over long videos because they forget parts of the earlier frames due to memory limits during generation. The authors found this happens because models are trained on short sequences but have to generate long sequences at test time, causing important context to be lost. They propose a way to gently reduce how much older frames influence prediction instead of abruptly cutting them off, which helps the model remember longer histories without losing motion continuity. Their method works without changing training or needing extra memory, leading to better long video generations.

What this means in practice

Authors

Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen

Abstract

Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( Δ, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.