Deep reinforcement learning improves precision in risky motion tasks

Learning High-Risk High-Precision Motion Control

Machine LearningRobotics

Summary

Some tasks need very precise moves where one mistake cannot be fixed later, like playing a tricky billiard shot. Most AI methods do okay when they can correct mistakes later, but struggle with these high-risk, high-precision tasks. The authors made a new AI method called SCOOT that learns from only the best moves, switches strategies depending on the situation, and tries many approaches before focusing on the best ones. They tested this on billiard shots and showed it could find very precise and varied ways to hit the balls.

What this means in practice

  • For robotics engineers: Enable robots to safely perform precise, irreversible tasks by learning from elite successful strategies under high-risk conditions.
  • For game developers: Create realistic AI-controlled characters that perform complex, precise movements without allowed mistakes, improving game physics and challenge.

Authors

Nam Hee Kim, Markus Kirjonen, Perttu Hämäläinen

Abstract

Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using elite samples, allowing the policy to better latch on to the rare high-reward action samples; 2) Utilizing a mixture-of-experts (MoE) policy, to allow switching between reward landscape modes depending on the state; 3) Adding a distance regularization term and a learning curriculum to encourage exploring diverse strategies before adapting to the most advantageous samples. We showcase our features' performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple shot strategies for a given ball configuration.