VicEdit: Learning to Edit Videos from Visual In-Context Examples
2026-08-17 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors address the challenge that text instructions alone have limits in describing detailed changes in video editing. They introduce a new method called Visual In-context Editing, which uses visual examples like images or videos alongside text to guide edits more precisely. To support this, they created a large dataset named VicEdit-400K with diverse video editing tasks. They also developed a model called VicEdit that smartly combines information from both visuals and text to improve editing quality. Tests show their approach works better than previous methods for various video editing tasks.
video editingvisual in-context editingtextual instructionsmulti-modal guidancedatasetsemantic tokensmodality-adaptivedual-context injectionvisual fidelitysemantic consistency
Authors
Yuji Wang, Teng Hu, Yuheng Chen, Ran Yi, Han Feng, Weijian Cao, Chengjie Wang, Lizhuang Ma, Jiangning Zhang
Abstract
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.