Speech drives realistic 3D facial animation with visible lip movements

Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation

GraphicsMachine Learning

Summary

It is hard to create 3D facial animations that match spoken words perfectly because many mouth shapes can produce the same sounds. The authors built a new system that understands mouth movements in three directions—spreading, opening, and sticking out—to better match speech. Their method uses a special memory to link sounds to these movements and combines them in a way that keeps the 3D face model smooth. Tests showed this makes lip animations more accurate and natural looking than previous methods.

What this means in practice

  • For game developers: Generate accurate lip sync and realistic facial expressions from speech audio to improve character immersion in games.$Commercial implications: Enables selling games with high-quality speech-driven facial animation enhancing player experience and realism.
  • For virtual assistant designers: Create speech-driven 3D avatars with more natural mouth and lip movements to improve user interaction and trustworthiness.

Authors

Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim

Abstract

Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.