Llm multi-agent systems improve collaboration by tracking proposal changes

MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

Multiagent Systems

Summary

Figuring out how groups of AI agents work together to solve problems is tricky because most tests only look at the final answer, not how they got there. The authors created MASTraceBench, a way to watch and measure how agents share and improve their ideas during teamwork. They found that the best final answers usually come from the strongest early idea and that weaker ideas often get boosted while strong ones rarely get better. To fix this, they made a new method called CLEARS that checks smaller parts of proposals between agents, helping keep or improve the best ideas more often.

What this means in practice

  • For ai system developers: Diagnose and improve multi-agent collaboration by analyzing how proposals evolve and are combined to enhance final decisions.
  • For software engineering teams: Guide design of AI systems that coordinate multiple agents by using claim-level evaluation to maintain or improve the strongest ideas during collaboration.

Authors

Yapeng Li, Songze Li, Shuang Yu, Jing Yu, Zhixin Liu, Liqiang Wen, Tonghua Su

Abstract

LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.