Robot adapts shared control to human behavior for better cooperation
Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation
Robotics
Summary
Robots and humans sometimes work together to control machines, but humans don’t always act perfectly logically. The authors developed a way for a robot to guess what kind of thinking a human partner might have during control, using patterns of past actions. Instead of trusting just one guess, the robot considers a range of possible human behaviors to decide how to assist. This approach helps the robot provide better support, shown in tests where the robot’s actions matched the human’s style more closely and reduced errors.
What this means in practice
- •For robotics engineers: Improve robot assistance in human-robot shared tasks by adapting to variable human decision styles to reduce errors in nonlinear control systems.
- •For automation system designers: Design adaptive controllers that incorporate uncertain human behavior models to enhance cooperation in complex manipulation tasks.
Authors
Henry Ascencio Trejo, Roel Pieters, Gokhan Alcan
Abstract
This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.