Progressively Learning Heterogeneous Skills in a Unified Latent Space
2026-08-24 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors introduce HetSkills, a system that teaches a computer how to control a character's movements by learning different skills all in one shared space. They start by teaching the character to imitate motions well, then add the ability to follow language instructions without taking shortcuts, using special techniques to connect words to movements. This approach helps the character keep natural-looking actions while learning more abilities over time. Tests show HetSkills works well for various tasks like tracking, generating motions from text, and finishing motions, even in hard situations.
latent spacephysics-based character controlmotion trackingtext-to-motion generationmotion decoderlanguage semanticstask-guidance modulemotion completionskill integrationlong-horizon tasks
Authors
Yue-Yi Zhang, Ming Gong, Linpu He, Wei-Shi Zheng, Zhilin Zhao
Abstract
We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-based character control. The core idea is to treat this latent space as a shared executable interface, enabling seamless integration of skills learned from diverse data sources, supervision forms, and tasks. HetSkills begins by learning a tracking skill that establishes a strong foundation in motion control and creates a shared motion decoder, which can be reused across tasks without the need for retraining or separate controllers. To prevent the text-to-motion skill from exploiting shortcut pathways instead of learning language semantics, we introduce motion intuition distillation to ground text-to-motion generation in language semantics and a task-guidance module that dynamically adjusts actions based on high-level language instructions. This enables HetSkills to preserve natural motion while continuously expanding its skill repertoire, making it highly adaptable for long-horizon tasks. Experimental results demonstrate the effectiveness in motion tracking, text-to-motion generation, motion completion, and downstream task adaptation, achieving impressive success rates even under challenging conditions.