Benchmarking audio video generation with multiple reference inputs

ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts

Computer Vision and Pattern Recognition

Summary

Creating videos with sound using multiple reference clips is tricky because the system must understand and combine different elements correctly. The authors created ORAV Bench, a set of tests with many examples where multiple references guide the generation of new audio-video clips. They also designed a method to evaluate how well the generated clips match the intended content by comparing pairs of outputs, closely aligning with human judgments. Their tests show that current systems can struggle to avoid copying unintended parts from source clips, revealing areas for improvement.

What this means in practice

  • For video content creators: Evaluate and improve tools that generate videos with sound using multiple reference clips to ensure accurate blending of features.
  • For machine learning engineers: Test and diagnose audio-video generation models’ ability to follow multiple input references, guiding targeted improvements in generative AI systems.

Authors

Jiacheng Hua, Xiaokun Feng, Jiaqi Hua, Chang Liu, Biao Wang, Miao Liu

Abstract

Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-10 references, 9 semantic roles, and 30 role compositions. Instructions specify the relationships among references; the media supply the identities, dynamics, and audio characteristics to be realized. To evaluate these open-ended outputs, we develop a reference-aware pairwise protocol that prepares visual and auditory evidence, compares the intended contribution of each reference, and checks the overall verdict in both presentation orders. On held-out instances, it achieves 86.08% effective agreement with human judgments. Across 5 frontier systems, overall rankings conceal distinct strengths across reference compositions. A recurring failure is to reproduce unintended source content in place of the requested result, despite closely resembling a reference. Reproducible pointwise diagnostics of quality, reference affinity, and speech reveal distinct dimensions of model behavior. ORAV thus offers a benchmark for tracking progress toward controllable, compositional, and reference-faithful audio-video generation.