Humanoid robot control improves with discrete action token model

Humanoid Loco-Manipulation With Discrete VLA Model

Robotics

Summary

Controlling humanoid robots is difficult because they have many body parts that move in different ways. The authors developed a new model called Holo-M that breaks down these complex movements into smaller parts, each represented as special tokens. This approach uses a language model extended to understand these action tokens, allowing the robot to learn from different types of movement data. This helps the humanoid robot perform walking and manipulating tasks more successfully and quickly than previous methods.

What this means in practice

  • For robotics engineers: Control humanoid robots performing complex walking and manipulation tasks using a unified discrete token model for real-time operation.
  • For virtual reality developers: Generate humanoid avatar movements from mixed real and simulated motion data for more realistic loco-manipulation in VR environments.

Authors

Wenxin Shao, Siqi Chai, Kun Li, Kerou Zhang, Xinzhou Jiang, Wei Xu, Qiang Liu

Abstract

Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.