Voice agents struggle to complete professional workflows reliably

APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction

Artificial Intelligence

Summary

Voice agents that can talk and listen at the same time are good at having conversations but often fail to finish complex tasks correctly. The authors created APEX-Voice, a test with 120 real-world professional tasks like filling forms and negotiating. They found that even the best voice agents succeed less than 25% of the time in completing these workflows properly. Most failures happen because the agents can’t keep track of ongoing tasks well, especially when they need to search for information or fix mistakes mid-conversation. This work shows that being able to chat does not mean a voice agent can reliably do professional work.

What this means in practice

Authors

Puneet Mathur, Dinesh Manocha

Abstract

Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.