Selecting models for multi-agent games without perfect info on strategies

Rank Without an Oracle: Deviation-Aware Interaction-Rank Selection from Offline Multi-Agent Logs

Multiagent SystemsComputer Science and Game Theory

Summary

When predicting how players behave in games where multiple agents interact, it’s tricky because the data comes from past play but we want to understand new, changed strategies. The authors present a method called Selective Interaction-Rank Validation (SIRV) to better choose models that remain accurate even when players change their strategy unilaterally. Their approach tests each candidate model carefully across many possible scenarios to avoid picking one that works only on the original data but misrepresents incentives. They also provide mathematical guarantees on how well the chosen model predicts new situations and demonstrate improvements in prediction accuracy over previous methods. This helps in understanding strategic behavior without needing perfect knowledge of all possible moves.

offline multi-agent learningpayoff modelsstrategy deviationmodel selectionvalidation splitcorrelated equilibriumrisk boundsempirical Bernstein boundsgame theorynon-identifiability

Authors

Xiangwu Wang, Chengwei Cao, Hongyuan Tang

Abstract

Offline multi-agent payoff models are estimated under a logging distribution but used on distributions induced by learned solutions and unilateral deviations. Standard held-out loss can therefore favor an interaction class that predicts logged play well while distorting strategic incentives. We introduce Selective Interaction-Rank Validation (SIRV) for finite games with known logging distributions. A training split fits nested payoff models and constructs a common union of all candidate deployment and unilateral-replacement distributions; an independent calibration split evaluates every candidate on this same union. SIRV returns the smallest rank whose simultaneous upper worst-target risk is within tolerance of the best upper score, and abstains when a declared target is unsupported or too imprecisely estimated. A common coverage event yields a finite-candidate target-risk bound and a candidate-specific coarse correlated equilibrium (CCE) gap certificate. We also isolate an exact two-point off-support non-identifiability result. In a controlled factorial study with 2,048 independent games per family, empirical-Bernstein bounds reduce the median CCE-gap certificate by 42.5% relative to Hoeffding bounds on common returns, with a 1.36-point reduction in supported return. Under paired rank misspecification and in a separately generated congestion family, the SIRV-EB fallback rule lowers mean true candidate-selection CCE regret relative to ID-Mean, while retaining game-level losses. Across 384 games at $N=3,5,8$, ID-Mean-relative mean CCE-regret effects stay positive while certified return falls sharply under weak coverage. These results separate certifiable model selection from universal strategic improvement.