AgentWorld benchmarks long multi-agent collaboration in games

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Multiagent Systems

Summary

Many current tests for AI agents focus on short or competitive activities and don’t really measure how well multiple agents work together over time. The authors created AgentWorld, which tests teams of AI agents collaborating in a complex online game for many steps and with different roles. They also made a new way to measure how much each agent’s actions contribute to the team's success. When tested on several advanced AI models, none did better than about half the tasks, often failing because agents didn’t communicate well or keep plans. The benchmark and tools are open-source for others to use.

What this means in practice

  • For game developers: Evaluate AI agents’ ability to collaborate over many moves in complex virtual environments for more realistic team behaviors.
  • For chatbot developers: Test and improve AI systems that require multiple specialized bots to coordinate tasks by communicating and planning jointly.

Authors

Raphael Shu, Yusen Zhang, Young Min Cho, Jin Mo Yang, Yuan Yuan, Wenliang Zheng, Sharath Chandra Guntuku, Lyle Ungar, Zhou Yu, Rui Zhang

Abstract

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.