Dynamic depth speech enhancers match static models in quality and speed
Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
Machine LearningSound
Summary
On-device speech enhancers, like those in hearing aids or earbuds, often run fixed-speed models. The authors trained one deep learning model to work well at multiple depths, allowing it to run faster or slower dynamically without losing quality. They tested these models on small microcontrollers and found that dynamically choosing the model depth per audio frame does not harm performance or add much delay. This means devices can adjust computation on the fly, saving power while maintaining sound quality.
What this means in practice
- •For embedded device engineers: Use dynamic depth speech enhancement to reduce compute load without sacrificing audio quality on constrained devices.
- •For audio hardware manufacturers: Develop hearing aids and earbuds that adapt compute dynamically per audio frame, improving battery life and responsiveness.$Commercial implications: This enables more efficient speech enhancement products targeting consumer wearables with limited compute capacity.
Authors
Clément Laroche, Riccardo Miccini
Abstract
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 $μ$s per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.