Turn-taking in dialogue models improves by using speaker intent timing
Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models
Computation and LanguageArtificial IntelligenceHuman-Computer InteractionSound
Summary
Current methods to judge when one person should speak after another in conversations are too simple, only allowing quick replies or silence. The authors show this isn’t enough because people sometimes wait or talk over each other depending on their intentions. They created a new benchmark called TACT with real dialogue data and detailed information about what speakers intend to do. This benchmark scores turn-taking by comparing timing patterns to those of humans, improving agreement with human judgments. Their tests show models can get better at understanding conversation flow beyond simple rules.
What this means in practice
- •For conversational ai developers: Improve training and evaluation of dialogue systems by considering speaker intent for natural timing of turn-taking.
- •For virtual assistant designers: Develop assistants that better handle interruptions and pauses based on user intent, making conversations more natural.
Authors
Kian Shamsaie, Iman Modarressi
Abstract
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.