LYRIC enables realistic whole-body object interaction from language commands
LYRIC: Language-Driven Physics-Based Character Control for Contact-Rich Whole-Body Object Interaction
RoboticsGraphics
Summary
Controlling virtual characters to interact with objects in realistic ways is hard, especially when instructions are given in plain language. The authors developed LYRIC, a system that lets simulated characters perform complex whole-body tasks using only natural language instructions and a goal for the object they interact with. LYRIC uses a two-part controller to plan short moves and then generate detailed body motions and contacts. It performs better on standard tests than past methods and can handle new objects without extra training.
What this means in practice
- •For game developers: Create interactive game characters that follow natural language commands to perform realistic whole-body object manipulations.$Commercial implications: Enables development of more immersive games with responsive, physics-based character control from language, appealing to gaming companies and studios.
- •For robotics simulation teams: Simulate complex contact-rich interactions between humanoid robots and objects using language instructions to test control strategies.
Authors
Zeyu Han, Zichong Meng, Julian Tanke, Minami Matsumoto, Sergey Bashkirov, Yingruo Fan, Selim Engin, Dongseok Shim, Takashi Shibuya, Yuki Mitsufuji, Huaizu Jiang
Abstract
We present LYRIC, a generative flow-matching controller for language-driven physics-based contact-rich interaction control, that enables simulated characters to perform contact-rich whole-body object interactions from a free-form language instruction and a sparse terminal object goal. To obtain reliable expert trajectories from imperfect motion-capture references, a single tracking policy is trained using geometry-conditioned interaction rewards and relaxed reference tracking near hand-object contact. To guide interaction progress without prescribing a full-body kinematic reference, we factorize the controller into a task-level planner that predicts short-horizon object and humanoid-root trajectories, and an action generator that resolves whole-body motion and contacts in closed loop. After behavior cloning, we freeze the planner and post-tune the action generator on policy using the planner's predictions as stable supervision for intermediate task progression. In a controlled OMOMO evaluation, our tracker achieves 64.3% success compared with 53.2% for an InterMimic reimplementation, while a unified policy achieves 76.5% on the full OMOMO dataset. On the held-out split, LYRIC achieves 90.3% task success, compared with 74.2% for the strongest matched kinematic-planner baseline, with better semantic alignment and motion quality. Without retraining, the controller also supports test-time object-waypoint guidance. Qualitative results further demonstrate robust, natural contact-rich interactions and zero-shot transfer to novel object shapes. The webpage is available at https://neu-vi.github.io/LYRIC/