Summary
Learning to copy experts by watching their actions can be tricky, especially when using adversarial methods that try to spot differences between the learner and expert. The authors study how adding certain kinds of regularization, which gently guide the learning process, can help the learner copy experts faster and more reliably with less data. They design a new learning algorithm that balances penalties on both the policy and reward, proving it learns quickly as it gathers more expert examples and tries new actions. This helps explain why regularization works well in practice and gives guarantees on how fast learning happens even when the expert’s behavior is noisy.
What this means in practice
- •For robotics engineers: Improve robot training efficiency by applying regularized imitation learning that quickly matches expert demonstrations with fewer trials.
- •For autonomous vehicle developers: Enhance driver behavior imitation in self-driving cars through algorithms that balance policy and reward penalties for faster learning from expert data.
Authors
Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
Abstract
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for jointly regularized AIL in finite-horizon Markov decision processes with general function approximation. Our model-free algorithm, Dually Regularized AIL, combines KL policy regularization with a quadratic reward penalty weighted by expert and learner occupancies. With K online episodes and N expert trajectories, we prove a $\widetilde{O}\left(\frac{1}{K}+\frac{1}{N}\right)$ bound on the regularized imitation gap for fixed regularization parameters. Our analysis combines an online mirror descent construction for general convex reward classes to control estimation error from finite expert data and stochastic learner feedback, with a sharp analysis of optimistic KL-regularized policy learning. To the best of our knowledge, Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}ε\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts. These results provide a rigorous characterization of the complementary statistical benefits of reward and policy regularization in AIL.