Driving scores lose meaning when used for direct optimization

When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving

Computer Vision and Pattern RecognitionRobotics

Summary

Driving test scores are often used to see how well an autonomous car drives. But when those scores become the exact goal to improve, the link between a better score and safer driving can break down. The authors show that some changes in driving don’t affect the score much, and that improving the score can fail when the way the car follows instructions changes. This means score improvements may not always reflect real driving improvements unless we consider both how the score is measured and how the car drives.

What this means in practice

  • For autonomous vehicle developers: Design evaluation and training processes that align score improvements with actual driving behavior changes for safer autonomous cars.
  • For robotics control engineers: Avoid relying solely on optimized benchmark scores when tuning control interfaces to ensure improvements transfer to real-world driving.

Authors

Morui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi, Qing Yang

Abstract

Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring process responds to changes in driving behavior and whether the resulting gains persist under repeated execution and replanning. We decompose the process into execution, measurement, subscore mapping, and aggregation. Controlled interventions reveal substantial behavioral changes that receive little score response because distinctions are omitted, thresholded, or attenuated between requested and executed motion. Closed-loop comparisons further show that optimization gains can reverse when the execution interface changes, demonstrating their dependence on how requests are executed and returned as feedback. Together, these findings connect the behavioral distinctions preserved by a metric to the conditions under which its gains transfer. Metric validity under optimization therefore requires examining both what the scoring process measures and how the optimized behavior is executed.