Multi-agent communication improves solving complex tasks and compression

Scaling Discovery through Test-Time Communication

Machine LearningArtificial IntelligenceComputation and Language

Summary

Solving hard problems often gets easier when people share what they discover, instead of working alone. The authors show that when multiple AI agents talk to each other during their work, they solve tough puzzles better than many agents working independently. This teamwork leads to breakthroughs that a single agent could not achieve. They tested their idea on challenge problems like puzzle packing and compressing handwritten digit classifiers, finding better results than the best known solutions. However, this advantage depends on enough computing power and clear feedback during problem solving.

What this means in practice

  • For ai system builders: Build cooperative AI agents that share discoveries during problem solving to improve success rates on complex tasks under sufficient compute.
  • For data compression engineers: Create smaller, high-accuracy classifiers by coordinating multiple compression agents communicating at test time.

Authors

Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos

Abstract

Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.