Behavioral foundation models improve diverse robotic skill discovery
Behavioral Foundation Models for Quality Diversity
Machine LearningArtificial Intelligence
Summary
Finding many different ways for robots to do tasks well is a hard problem, especially when the tasks give little feedback or are tricky. The authors studied special models called Behavioral Foundation Models (BFMs), which learn a compact way to describe robot behaviors. They showed that searching for new robot actions inside this compact space works better than searching in the usual big space of robot controls. This method finds many good and varied robot behaviors, especially when simple approaches fail. The work suggests BFMs can help robots learn many useful skills from stored data, beyond just one task.
What this means in practice
- •For robotics engineers: Generate a wide variety of efficient robot behaviors from pre-trained models without training new policies from scratch.
- •For autonomous vehicle developers: Improve navigation strategies in sparse or deceptive environments by optimizing policies in a compact behavior space.
Authors
Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud
Abstract
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all tested parameter-space QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.