Multi-agent reinforcement learning gains interpretable concepts for cooperation

Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning

Artificial Intelligence

Summary

Cooperation is hard when each agent sees only part of the whole picture. The authors created a system where each agent learns simple meaningful concepts from its actions and observations to explain its choices. These concepts come together to form a bigger picture that helps agents work together better and show how decisions are made. Their method also encourages agents to explore new ideas by rewarding them when they predict new concepts. Tests show this approach performs well and helps understand agent behavior.

What this means in practice

  • For robotics engineers: Implement interpretable multi-agent decision-making systems where each robot explains its choices during cooperative tasks.
  • For game ai developers: Build game agents that show understandable cooperative strategies, improving debugging and game design.

Authors

Yijie Sun, Sanquan Sun, Yanda Zhu, Yuanyang Zhu, Yaohua Hu, Chunlin Chen

Abstract

Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.