Joint audio video generation improves fidelity and sync with new RL method
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Computer Vision and Pattern RecognitionSound
Summary
Generating videos with matching sounds is hard because models struggle to make each part realistic and well-timed together. The researchers created AV-GRPO, a new learning approach that separates audio and video training to give clearer feedback and better timing. They also made 5DAV, a special dataset to help train models more effectively. Tests showed their method produces better quality videos and sounds that match well with the text descriptions and each other.
What this means in practice
- •For multimedia content creators: Create videos with synchronized audio and visual elements that align better with script input, improving quality for creative productions.$Commercial implications: This enables production studios and content platforms to offer higher fidelity audio-video experiences aligned with user or automatic text prompts.
- •For multimodal ai system developers: Use AV-GRPO’s training techniques to improve multimodal model training efficiency and synchronization evaluation during joint audio-video generation.
Authors
Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao
Abstract
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computationally expensive given their divergent dynamics. Moreover, synchronization evaluation difficulty depends on paired samples, preventing fair reward comparisons. We propose AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset. AV-GRPO includes three key modules: (1) modality-anchored rollouts to disentangle learning signals and stabilize difficulty; (2) trajectory-locked frozen-tower optimization to reduce cost and reassign credit; (3) adaptive objectives and perturbation strengths tailored to modality-specific dynamics. This converts coupled multimodal preference learning into unimodal subproblems for precise reward attribution and better synchronization. Our 5DAV dataset decouples samples across five dimensions for systematic training. Experiments on JavisBench and VABench demonstrate AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning. Ablations confirm our designs. Code and data: https://github.com/zhiyuxu03/AV-GRPO