AI Alignment through a Game-theoretic Lens: A Survey

Artificial IntelligenceComputation and LanguageComputer Science and Game Theory

Summary

The authors review how aligning AI systems, like large language models, with complex human values is tricky because preferences can change, conflict, and depend on social context. They look at this problem using game theory, a way to study strategic interactions among different players. Their survey organizes recent work by focusing on challenges like varied preferences, deciding what to prioritize, and how things change over time. They explain where game theory helps improve AI alignment and where it still falls short in making AI systems safe and reliable.

Authors

Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang

Abstract

As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.