Voice agents struggle to complete professional workflows reliably
APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
Artificial Intelligence
Summary
Voice agents that can talk and listen at the same time are good at having conversations but often fail to finish complex tasks correctly. The authors created APEX-Voice, a test with 120 real-world professional tasks like filling forms and negotiating. They found that even the best voice agents succeed less than 25% of the time in completing these workflows properly. Most failures happen because the agents can’t keep track of ongoing tasks well, especially when they need to search for information or fix mistakes mid-conversation. This work shows that being able to chat does not mean a voice agent can reliably do professional work.
What this means in practice
- •For customer support teams: Evaluate and improve voice assistants that manage multi-step customer service tasks to reduce failures in task completion.
- •For office automation developers: Develop voice-enabled tools for business workflows that track task state and handle corrections effectively.
Authors
Puneet Mathur, Dinesh Manocha
Abstract
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes such as form completion, corporate negotiation, coordination, consulting, and interviewing. Each workflow executes in a stateful Voice Workbench environment with task-specific knowledge, typed tools, gold-annotated final work artifact, authorization constraints, and a user simulation policy backed by validated, pre-compiled speech realizations. We evaluate both artifact field accuracy and end-to-end workflow success, which requires the correct terminal state, valid process, completed actions, and a valid final artifact. Across five frontier real-time voice agents-GPT-Live-1, Gemini-3.8-Live, Grok-Voice-Think-2.0, Step-Audio3, and GPT-realtime-2.1, none exceeds 25% Pass@1, and the best Reliable@3 is only 10.8%. Moreover, stateful coordination is the dominant failure point across systems, while success decreases further on workflows requiring greater knowledge retrieval and mid-speech corrections. Overall, APEX-Voice is the first benchmark for evaluating whether voice agents can translate conversational competence into dependable professional work.