TEMA improves tracking and timing answers in multi-audio dialogs

TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs

Sound

Summary

Answering questions about events happening in multiple audio clips over several turns is tough because you have to remember what happened before and when. The authors created TEMA, a new system that links hearing events with clear evidence to answer these tricky questions. They built a large dataset and tests to train and check TEMA, which helps spot events and compare them across different audio sources more accurately. They also improved it by using a special training method that focuses on finding evidence first.

What this means in practice

  • For voice assistant developers: Enhance voice assistants to answer detailed timeline questions spanning multiple audio recordings with evidence-backed responses.$Commercial implications: Enables voice assistants to handle complex, multi-turn audio queries for commercial products requiring accurate temporal event understanding.
  • For audio forensic analysts: Assist forensic teams in pinpointing and comparing event timings across multiple audio files during investigative reviews.

Authors

Kaidi Yang, Hualei Wang, Zhaohui Wang, Chenxuan Wang, Hong Liu, Xiangdong Wang

Abstract

Multi-turn, multi-audio temporal question answering requires models to track target events across follow-up questions, recording switches, and historical references, recovering complete instances and their boundaries for temporal calculation and comparison. We propose TEMA, which connects event perception with evidence-based answering through Route, specifying the audio scope, and Span, describing all relevant intervals as conditional audio captions. We construct TEMA-Dialog with 40,704 dialogs and per-turn evidence and answer supervision, and TEMA-Bench for joint evaluation of evidence and final answers. Training combines temporal grounding initialization, full-dialog supervised fine-tuning, and completeness-first Span-only GRPO. Experiments on Qwen2.5-Omni and AF-Next show improved temporal question answering, particularly event localization and cross-audio comparison. Reinforcement learning applied solely to evidence further improves interval recovery and answer accuracy.