Speaker verification system explains voice decisions with acoustic details
CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification
Computation and LanguageSound
Summary
Current speaker verification systems can tell if two voices are from the same person but do not explain how they made that decision. The authors created CoLMbo-SV, which combines a voice recognition model with a language model that produces detailed, understandable reports that refer to specific acoustic measurements. These reports help users inspect and understand the evidence behind verification results without limiting the system’s accuracy. They also introduced VoxReason, a dataset with paired voice recordings, measurements, and verified comparison reports to train and evaluate such models.
What this means in practice
- •For security system developers: Build voice authentication tools that provide clear explanations of identity verification results with grounded acoustic evidence.$Commercial implications: Enables sale of transparent biometric security products that increase user trust by explaining voice verification decisions.
- •For forensic audio analysts: Use detailed verbal and numerical reports to better understand voice comparison outcomes in legal or investigative contexts.
Authors
Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh, Bhiksha Raj
Abstract
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.