IndicFDB benchmarks full-duplex voice agents in ten Indian languages

IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages

Computation and Language

Summary

Handling natural conversation in voice agents—like knowing when to pause, interrupt, or respond—is tricky, especially in many Indian languages. The authors created IndicFDB, a big collection of Indian language voice samples to test how well voice agents manage these real-time conversations. They used smart methods to find conversational events and judge the agents’ timing and responses without relying on detailed text alignments. Their tests showed trade-offs between speed and handling conversation smoothly across languages.

What this means in practice

Authors

Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh, Manmeet Kaur, Sagar Jain, Hanuman Sidh, Pranav Sharma, Aditya Singh, Aaditya Pareek, Manas Dhir, Adi Margolin, Niket Agarwal, Bryan Catanzaro

Abstract

Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.