Edit VAR improves text guided video editing with better source and motion control
Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing
Computer Vision and Pattern Recognition
Summary
Editing videos by changing their content based on text is tricky because you need to keep parts of the video unchanged and consistent over time. The authors introduce Edit-VAR, a method that does this without needing extra training or reversing the video editing process. It smartly replaces parts of video data based on text, keeping unedited parts intact and making sure moving details stay smooth. Tests show it works better than existing methods in making edits that look good, keep the original video’s feel, and run efficiently.
What this means in practice
- •For video editors: Automatically modify video content based on text prompts while preserving unchanged video parts and ensuring smooth motion.$Commercial implications: Enables commercial video editing software to offer precise, training-free text-based editing features that preserve original video quality.
- •For visual effects artists: Enhance video post-production by editing specific visual elements in a scene without retraining or complex inversion steps.
Authors
Chongbo Zhao, Jiangming Wang, Xilai Wang, Xinyu Wang, Jingyi Tang, Chunjie Hao, Pengjie Song, Yue Ma
Abstract
Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.