Benchmark reveals how communication shapes multi-agent system performance
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
Machine Learning
Summary
Coordinating multiple AI agents involves graphs that dictate how they share information and split tasks, but it's hard to tell which part of this setup affects their success. The authors created OpenMAS-GCom, a test environment that changes one piece of the system at a time to pinpoint how communication links, roles, and information flows impact performance. They tested many configurations on diverse tasks, discovering that losing specialist roles hurts more than losing critics, and that wrong messages or worker failures degrade performance differently. This benchmark helps understand what really matters in multi-agent communication setups.
What this means in practice
- •For ai system developers: Diagnose which communication structures and roles most impact multi-agent system accuracy under fixed resources and tasks.
- •For software test engineers: Test system robustness by simulating agent removal or message corruption to find critical points in multi-agent communication.
Authors
Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li
Abstract
Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.