Human-tci improves text based retrieval of complex human motions

HUMAN-TCI: Hierarchical Multi-Stream Motion-Aware Network with Torso-Centered Interaction for Text-to-Motion Retrieval

Computer Vision and Pattern Recognition

Summary

Finding the right human movement sequences from text descriptions is hard because people can describe many actions at once and body parts move in connected ways. The authors created HUMAN-TCI, a model that looks at the upper body, lower body, and torso separately but also considers how torso movements influence the rest. This lets the model better understand complex actions and find matching motions even when sentences describe overlapping or multiple movements. Their approach is simpler and faster than previous methods while being good at handling mixed or longer action descriptions.

What this means in practice

  • For video game developers: Retrieve accurate human animation clips from natural language commands to streamline character motion design workflows.
  • For virtual reality developers: Improve selection of realistic human motions to animate avatars based on complex spoken or written instructions.

Authors

Muhammad Islam, Euijoon Ahn, Usman Naseem, Tao Huang

Abstract

Accurate retrieval of human motions is a crucial first step in text-guided human motion modeling and synthesis, as it selects semantically relevant sequences from large datasets and provides grounded references for downstream tasks. Retrieving motions from natural language descriptions remains challenging because sentences can describe multiple actions, overlapping movements, and intricate dependencies between body parts. Existing methods often focus on simple, single-action descriptions and typically process body parts independently or by merely concatenating features, without explicitly modeling how torso movements influence other parts. In addition, their processing pipelines often rely on computationally heavy models, introducing considerable overhead, particularly when modeling longer or more complex motion sequences. This limits learning discriminative motion-pattern representations, reducing retrieval accuracy, interpretability, and efficiency in practical applications. To address these limitations, we propose HUMAN-TCI, a Hierarchical Multi-Stream Motion-Aware Network for text-guided human motion retrieval. HUMAN-TCI employs a three-stream architecture that separately models upper-body, lower-body, and torso motions while explicitly capturing their interactions, allowing torso-related movements to influence the positioning and dynamics of other body parts. By incorporating tailored torso attention, our model effectively recognizes complex human motion patterns, captures fine-grained motion relationships and handles complex multi-action descriptions. Our framework supports retrieval for both simple, single-action sentences and long, compositional descriptions containing sequential or overlapping actions without relying on complex models.