AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

2026-07-09Machine Learning

Machine LearningArtificial IntelligenceComputer Science and Game Theory
AI summary

The authors studied how well AlphaZero plays two different games, Connect Four and Chomp, comparing its usual version to a new one that uses extra help from an expert. They found that the standard AlphaZero plays very well but doesn’t always follow the perfect winning moves in these games. Adding a special training technique, called AlphaZero Auxiliary Loss (AZAL), helped it make fewer mistakes and get closer to perfect play, especially in Chomp. However, even with these improvements, perfect play was not fully achieved in either game.

AlphaZeroMonte Carlo Tree Search (MCTS)self-playConnect FourChomp gameoracle supervisiongame-theoretic valueGrundy numberpolicy supervisionperfect play
Authors
Brent Kong, Tejas Ram, Tony Yue Yu
Abstract
AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game with exact game-theoretic values, and Chomp, an impartial game whose optimal play is governed by Grundy-number structure. Under a unified self-play $+$ MCTS pipeline, we compare vanilla AlphaZero, a multi-frame variant (limited to Chomp), and an AlphaZero Auxiliary Loss (AZAL) that adds oracle-derived policy supervision. We find that vanilla AlphaZero achieves strong play across both domains but cannot preserve the exact trajectories required for optimal play: in Connect Four, it fails to maintain the optimal line of play, while in Chomp, it fails to consistently restore the $g=0$ invariant. On rectangular Chomp boards, multi-frame inputs alone do not remove this gap. Nevertheless, AZAL substantially improves oracle consistency across multi-seeded full-game traces and sampled-state evaluations. On Chomp, AZAL reaches perfect full-game oracle consistency on 10x11 and high but not complete consistency on 9x10; on Connect Four, AZAL improves oracle-match rate and delays the first oracle mistake, but does not reach perfect play.