Benchmark evaluates AI email agents on enterprise productivity tasks
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
Artificial Intelligence
Summary
Handling enterprise email involves tasks like retrieving information, managing schedules, and coordinating multiple steps accurately. The authors created EmailBench, a set of 206 email-related tasks designed to test how well AI agents can perform these real-world workflows using a synthetic email dataset. They tested eight AI configurations and found that even the best one completed only about a third of tasks correctly, showing that just completing API actions doesn’t mean the task was truly done. This benchmark helps developers measure and improve AI email assistants in a controlled, realistic setting.
What this means in practice
- •For enterprise software developers: Test and improve AI agents for managing complex email workflows in business environments using a standardized task set and evaluation protocol.
- •For ai tool integrators: Evaluate the effectiveness of language models combined with APIs for automating productivity tasks involving email and scheduling.
Authors
Mukul Singh, Mansi Uniyal, Devin Devlin, Wen Xie, Big Thadawasin, Ritam Dutt, Vivian Lai, Hyeonsu B. Kang
Abstract
Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.