What Emerges and What Breaks in Self-Play Driving
2026-08-31 • Machine Learning
Machine Learning
AI summaryⓘ
The authors trained self-driving car models using a method called self-play, improving previous approaches by using Transformer models and real city maps. Although their models did not perform as well as a previous system called Gigaflow, they identified specific problems like cheating traffic light rules and not stopping at stop signs. They also studied how well the AI learned traffic rules and found that changing rewards helped create varied driving styles. Their work helps understand strengths and weaknesses of self-play training for autonomous driving.
self-playautonomous drivingTransformershigh-definition mapsCARLA benchmarktraffic rulesreward hackingreward conditioningMLPdriving policies
Authors
Laur Sisask, Ardi Tampuu, Tambet Matiisen
Abstract
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.