Method converts multi-agent AI failures into trainable theory-of-mind tests
ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures
Multiagent Systems
Summary
Communication between AI agents doesn't always lead to success because agents can misunderstand each other's roles or intentions. The authors created a method called ToMAS that turns these misunderstandings into specific test items that measure how well an AI can reason about other agents' states. They showed that humans can reliably identify these items and built a way to use them in training, but their initial training experiment did not show improvement due to technical reasons. ToMAS sets up a useful process for future work to better train AI agents to work together.
What this means in practice
- •For multi-agent system developers: Identify coordination failures to create targeted training tasks that improve AI agents’ understanding of each other's roles and intentions.
- •For natural language processing engineers: Use converted partner-state reasoning items to design training benchmarks that test and enhance communication alignment between AI agents.
Tested on one dataset.
Authors
Muhammad Ashar Ishfaq, Glaucia Melo
Abstract
LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers' roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies four explicit convertibility criteria to diagnosed execution traces. A full conversion pass over 242 eligible non-AG2 training traces produced 39 CLEAN items. In an 18-trace reliability pilot, two annotators achieved 94.4% raw agreement and Cohen's kappa = 0.92. We then used the converted items as binary rewards in a small-scale GRPO feasibility experiment with Qwen2.5-1.5B. On a 28-item held-out Magentic GAIA diagnostic, every evaluated condition exceeded the ROUGE-L threshold on the same 2 of 28 items. Post-hoc adapter checks show why: under the learning rate used, the LoRA update remained numerically negligible (max abs Delta W about 7e-6), so all conditions decode identically to the untrained checkpoint. The experiment therefore does not show a training effect and cannot establish one; it reports an executable pipeline together with two limitations that any conclusive study must address: a provenance gap between the training and evaluation items, and lexical-overlap scoring. ToMAS provides a preliminary rubric and pipeline for converting diagnosed coordination failures into trainable partner-state reasoning items and identifies the requirements for a conclusive matched-domain evaluation.