Physical connection shapes communication limits in multi-agent teams

Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

Machine LearningMultiagent SystemsRobotics

Summary

When multiple agents try to work together with limited communication, understanding what messages should carry is hard. The authors found that physical interaction between agents influences how valuable communication is. For agents tightly linked physically, messages add little because the agents already sense enough. With no connection, simple messages solve the tasks, but when linked partially, learned communication falls short of the best possible. The study shows that the challenge is not how many bits can be sent, but how the agents are physically coupled.

What this means in practice

  • For robotics teams: Design communication strategies for robots that interact physically, focusing on when communication improves coordination versus when it is redundant.
  • For multi-agent software developers: Develop training methods for multi-agent systems where knowing physical interaction helps predict when learned communication will be effective or fail.

Authors

Mihir Chauhan, Aniket Bera

Abstract

Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.