A game theory for foundation models shows new paths to rational cooperation through similarity inference
2026-08-04 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors studied smart AI agents that make decisions together in situations where their choices affect each other, like social games. They found that unlike traditional game theory, which expects these agents to not cooperate, these AI agents tend to cooperate stably. This happens because these agents see themselves as part of the environment and consider that others might think like them, leading to mutual cooperation. The authors introduced a new theory called the 'embedded Bayesian agent' and a new solution concept, the 'embedded equilibrium,' to explain this behavior. This work helps understand how modern AI agents interact differently from classic models.
autonomous agentsfoundation modelsclassical game theorydecoupled agencyembedded agencyBayesian agentsocial dilemmasNash equilibriumembedded equilibriumoptimal planning
Authors
Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, João Sacramento, Blaise Agüera y Arcas
Abstract
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.