Sub-model memory convolutions cut power use in keyword spotting devices
Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
SoundMachine Learning
Summary
Devices that respond to spoken commands often struggle to work well on small gadgets like smartwatches because they need to be both accurate and very fast with limited battery and memory. The authors show a way to make these voice recognition systems more efficient by changing how the neural network processes voice input, reducing wasteful computation and saving power. Their method keeps the network easy to train and achieves good accuracy on a standard speech recognition test. This means voice commands could work better on tiny devices without draining batteries quickly.
What this means in practice
- •For wearable device engineers: Build voice command features for wearables with less battery drain and memory use while maintaining accuracy.
- •For smart home hardware teams: Implement low-power keyword spotting in resource-limited smart home gadgets to improve always-on voice control.
Authors
Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski
Abstract
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.