Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
2026-07-10 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors address the challenge of evaluating audio-visual generative models, which create synchronized sound and images, by focusing on how well these two match in terms of cause and effect. Existing methods mainly check if the timing aligns but fail when the generated content has errors or mismatches. To fix this, the authors created a dataset called SynthSync with human-ranked examples and developed a method using a large language model to turn these relative rankings into absolute scores. They also introduced a new optimization technique, resulting in a metric that better matches human judgments and helps benchmark future audio-visual generation models.
audio-visual generative modelscross-modal synchronizationstructural hallucinationspairwise human annotationsOmni-LLMlatent projectionReal-Valued Group Relative Policy Optimizationcausality in generative contenthuman preference alignmentbenchmarking audio-visual generation
Authors
Yijie Qian, Juncheng Wang, Chao Xu, Huihan Wang, Yuxiang Feng, Yang Liu, Baigui Sun, Yong Liu, Shujun Wang
Abstract
As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality in generated content. However, existing evaluation metrics presume structural correctness, reducing synchronization to mere temporal alignment. Consequently, they fail on generative outputs, especially when exhibiting structural hallucinations and asymmetric cross-modal relations, which currently \textbf{mandate expert human annotation to assess synchronization.} This dependency introduces a critical paradox: \emph{human evaluators rely on relative, reference-dependent comparisons, whereas automated metrics require reference-free, absolute scalars.} We resolve this paradox by proposing a framework that distills relative human perception into a continuous, globally consistent metric. First, we introduce SynthSync, a dataset of generative failures ranked via pairwise human annotations. Second, we adapt the Omni-LLM equipped with a continuous latent projection to translate relative human rankings into continuous absolute values. Third, we propose Real-Valued Group Relative Policy Optimization ($\mathbb{R}$-GRPO) to internalize the global causal structure of synchronization via listwise score distributions. Empirically, our metric achieves state-of-the-art human preference alignment. We leverage this estimator to establish a standardized benchmark, advancing AV-Gen assessment from low-level signal correlation to visually grounded causality.