Human motion created by combining spatial audio and text intent

MoSAT: Human Motion Generation from Spatial Audio and Textual Description

GraphicsComputer Vision and Pattern RecognitionRoboticsSound

Summary

People's movements are influenced by what they hear around them and what they want to do. This paper introduces a way to create human-like full-body motions using both sounds coming from specific directions and written descriptions of actions. The researchers also made a new dataset that matches motion data with spatial sounds and detailed text, helping their system learn better. Their method, called MoSAT, improves how natural and accurate the generated movements are by linking sounds, language, and motion smoothly. They also created special tools to check that the motions match the audio and text properly.

What this means in practice

  • For game developers: Create character animations that respond realistically to sounds and scripted actions for immersive gameplay experiences.$Commercial implications: Enables selling advanced game animation systems that dynamically generate context-aware human motions driven by sound and text input.
  • For virtual reality designers: Generate natural human motions in VR environments that react to directional sounds and user instructions to increase presence.

Authors

Shuyang Xu, Zhiyang Dou, Yiduo Hao, Zekun Li, Liang Pan, Jingbo Wang, Cheng Lin, Yuan Liu, Wenping Wang, Mingmin Zhao, Taku Komura

Abstract

Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial audio and natural language, a problem that has been largely overlooked in previous research. To support this task, We introduce STAM, a dataset of motion sequences paired with spatial audio and detailed textual annotations whose rich vocabulary affords precise and nuanced specification of human motions. We further introduce MoSAT, a latent flow-matching framework for full-body motion generation jointly conditioned on natural-language intent and directional spatial-audio cues through hierarchical cross-attention before generating motion. Such a hierarchical design enhances temporally coherent and semantically aligned motion sequences. We also develop tri-modal evaluators for comprehensive evaluation on this novel task. Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio's intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.