IndicFDB benchmarks full-duplex voice agents in ten Indian languages
IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages
Computation and Language
Summary
Handling natural conversation in voice agents—like knowing when to pause, interrupt, or respond—is tricky, especially in many Indian languages. The authors created IndicFDB, a big collection of Indian language voice samples to test how well voice agents manage these real-time conversations. They used smart methods to find conversational events and judge the agents’ timing and responses without relying on detailed text alignments. Their tests showed trade-offs between speed and handling conversation smoothly across languages.
What this means in practice
- •For voice assistant developers: Evaluate and improve the handling of conversational timing and interruptions across Indian languages in real-time voice assistants.
- •For customer service technology teams: Test the robustness and naturalness of multilingual voice agents for smoother customer interactions in Indian language markets.
Authors
Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh, Manmeet Kaur, Sagar Jain, Hanuman Sidh, Pranav Sharma, Aditya Singh, Aaditya Pareek, Manas Dhir, Adi Margolin, Niket Agarwal, Bryan Catanzaro
Abstract
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.