Gaussian blendshape distillation speeds up real-time avatar animation
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Computer Vision and Pattern RecognitionArtificial IntelligenceHuman-Computer InteractionMachine Learning
Summary
Animating detailed 3D avatars typically requires complex neural networks that slow down real-time performance. The authors found that the complex animations of pretrained Gaussian avatars can be expressed as a simpler mix of basic facial or body shapes called blendshapes. They created a method called GALA that uses a small network to quickly predict how much to mix these blendshapes, making animation much faster without retraining existing models. This approach works well across different avatar types, allowing smooth animations even on mobile devices.
What this means in practice
- •For game developers: Animate complex 3D characters efficiently in real-time games on mobile devices by replacing heavy neural models with fast blendshape-based approximation.$Commercial implications: Enables mobile games and apps to deliver detailed avatar animations smoothly, offering a practical product advantage in performance and battery life.
- •For virtual event platforms: Provide realistic avatar expressions and body movements during live video chats without requiring costly server-side processing.
Authors
Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
Abstract
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/