Sub-model memory convolutions cut power use in keyword spotting devices

Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device

SoundMachine Learning

Summary

Devices that respond to spoken commands often struggle to work well on small gadgets like smartwatches because they need to be both accurate and very fast with limited battery and memory. The authors show a way to make these voice recognition systems more efficient by changing how the neural network processes voice input, reducing wasteful computation and saving power. Their method keeps the network easy to train and achieves good accuracy on a standard speech recognition test. This means voice commands could work better on tiny devices without draining batteries quickly.

What this means in practice

Authors

Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski

Abstract

Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.