Video and audio generation and editing combined for better coherence
Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
Computer Vision and Pattern RecognitionMultimediaSound
Summary
Video and audio are usually created or edited separately, even though people experience them together. This paper looks at ways to create or change video and audio at the same time, making sure they stay matched up in time and meaning. The authors organize these methods into a single framework and list different categories of edits you can make when working with both audio and video. They also describe the data and measurements used to test these methods and highlight remaining challenges that need solving.
What this means in practice
- •For video production teams: Create or edit videos with matching audio automatically, ensuring synchronized and semantically consistent audiovisual content.
- •For sound design teams: Generate audio that aligns with given video scenes or edit sounds together with videos for coherent audiovisual storytelling.
A survey. It maps existing work.
Authors
Abhinav Sharma, Sai Karthik Navuluru, Wang Wei, Daksh Dangi, Xiangbo Gao, Li Li, Bo Ni, Vardhan Dongre, Junda Wu, Xiyang Hu, Jiuxiang Gu, Seunghyun Yoon, Tong Yu, Chien Van Nguyen, Mohamed Elmoghany, Nedim Lipka, Hoda Eldardiry, Hongjie Chen, Tyler Derr, Thien Huu Nguyen, Zhengzhong Tu, Nesreen K. Ahmed, Franck Dernoncourt, Ryan A. Rossi
Abstract
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.