Distributed Team Orchestration via Supervisor Networks: Convergence, Optimality, and Resilience

2026-08-10Multiagent Systems

Multiagent Systems
AI summary

The authors study games where teams of agents rely on supervisors for belief information, but this information might be wrong due to errors or deceptive behavior from some teams. They create a new method called DTOA that helps teams learn better beliefs and strategies, and prove it eventually leads to almost balanced team outcomes. When some teams act dishonestly, the authors modify DTOA to resist such attacks and can identify dishonest teams with some probability while still keeping the game near balanced. They also test their ideas through experiments showing how DTOA works compared to other methods.

zero-sum gamesteam-Nash equilibriumfictitious playbelief learningByzantine teamsdistributed algorithmsMarkov decision processgame theoryByzantine resilience
Authors
Juntian Zhu, Guanpu Chen, Tongtian Zhu, Miguel de Carvalho, Zhouwang Yang, Fengxiang He
Abstract
In this paper, we study zero-sum potential team games with a supervisor network, where agents rely on supervisor-provided belief information rather than accurate common beliefs. The main challenge is that such belief information can be inaccurate because of supervisors' belief-estimation errors and the misreporting of joint actions by Byzantine teams. We propose the distributed team-orchestrating algorithm (DTOA), which combines team fictitious play with supervisor-based distributed belief learning. We prove the convergence of supervisors' belief estimates and establish that the induced learning dynamics converge to a near team-Nash equilibrium (TNE) in terms of the team-Nash gap (TNG). In the Byzantine setting, we consider a misreporting attack model and develop a Byzantine-resilient DTOA. We further provide probabilistic guarantees for Byzantine-team identification and establish an asymptotic bound on the honest TNG. Numerical experiments illustrate the theoretical findings, compare DTOA with baseline learning methods, and evaluate its performance in a Markov decision process setting.