Papers for

multiagent system designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Optimistic hedge achieves lower regret in multiplayer games

A Logarithmic Regret Bound for Optimistic Hedge in General-Sum Games

Abstract: Can simple no-regret dynamics attain smaller regret in self-play than against arbitrary adversaries? In $n$-player general-sum games, Daskalakis et al. 2021 proved an $O(n\log d_i\log^4 T)$ individual regret bound for Optimistic Hedge, which improves upon the classical $O(\sqrt T)$ adversarial regret bound. In this work, we show that Optimistic Hedge with a constant step size can further achieve $O(\sqrt n\log d_i\log T)$ individual external regret under expected loss-vector feedback. The time-averaged play consequently enjoys a coarse correlated equilibrium gap $O(\sqrt n\log d\log T/T)$, where $d=\max_i d_i$. The improvement comes from a larger admissible step size $η=Θ(1/(\sqrt n\log T))$. Our analysis proves factorial bounds on high-order differences of probability-weighted pairwise loss gaps, then applies finite-difference interpolation in a fixed Euclidean norm. These estimates sharpen the analysis of Daskalakis et al. 2021 and yield a logarithmic regret bound.

Thu 17 SeptComputer Science and Game Theory
The gist
This work looks at how players can get better at making decisions in games where multiple people compete and cooperate. It shows that a specific learning method called Optimistic Hedge can learn faster and make fewer mistakes over time compared to previous approaches. The authors prove this by using a new mathematical analysis that allows the method to use a larger learning step. This improvement means players’ average strategies stabilize closer to an equilibrium faster than before.
Open → 2609.19677v1

Risk attitudes influence long-term outcomes in coordination games

Entropic Risk-Sensitive Evolutionary Learning and Equilibrium Selection in Coordination Games

Abstract: We study risk-sensitive evolutionary learning dynamics and their long-run equilibrium selection behaviors in coordination games. Agents' risk attitudes enter through the classical entropic risk measure, which evaluates opponent-induced payoff uncertainty and feeds into noisy best responses under two standard revision protocols: best response with mutations and logit choice. We first analyze $2\times 2$ coordination games in both single-population symmetric and two-population asymmetric settings. In the single-population setting, unlike the risk-neutral case where the dynamics are known to favor the risk-dominant equilibrium, we show that risk sensitivity can change the stochastically stable outcome: a greater risk-seeking attitude favors the payoff-dominant equilibrium, while a greater risk-averse attitude favors the maximin equilibrium. Thus, the population's risk attitude may act as a control knob for long-run equilibrium selection. In both population settings, we also identify a robust regime: any super-dominant equilibrium is stochastically stable for all risk attitudes, under both protocols, and across populations. We further extend the single-population analysis to symmetric $k$-action games, which include symmetric $k$-action coordination games as a special case, under risk-sensitive best response with mutations. In this setting, we show that, for sufficiently large populations, sufficiently risk-seeking agents uniquely select the strongly payoff-dominant equilibrium when it exists, whereas sufficiently risk-averse agents uniquely select the strongly maximin equilibrium when it exists. These results show that entropic risk sensitivity may serve as a systematic mechanism for steering equilibrium selection in evolutionary games, beyond the classical risk-neutral benchmark.

Tue 8 SeptComputer Science and Game TheoryMultiagent Systems
The gist
Sometimes people or groups have to choose together between different options that benefit them differently. This study looks at how being more or less worried about risk changes which shared choice is made over time. The authors show that when players are more risk-seeking, they tend to pick the option with the highest reward, but when they are more risk-averse, they prefer the safest option. These results suggest that how people feel about risk acts like a control that can steer which outcome becomes stable in repeated decisions.
Open → 2609.08677v1