Spiking neural networks can compute faster and more accurately with delay tuning

Latency and accuracy tradeoffs in Spiking Neural Networks

Machine Learning

Summary

Spiking neural networks (SNNs) are good at recognizing speech with low power, but people thought they were slower than other types of networks. The authors show that SNNs can actually finish faster by overlapping calculations across layers, but this can cause errors if spikes fire too soon. Surprisingly, waiting longer to fire doesn’t always help and can sometimes make things slower and less accurate. To solve this, the authors created Falcon, a method that finds the best balance between speed and accuracy by adjusting when each layer fires spikes.

What this means in practice

  • For mobile device engineers: Optimize low-power speech recognition systems by tuning spike timing to achieve faster response without losing accuracy.
  • For hardware accelerator designers: Design compute-in-memory architectures that leverage pipeline delay control to improve spiking neural network efficiency and speed.

Authors

Zhanglu Yan, Zixuan Zhu, Kaiwen Tang, Yuyang Cai, Qianhui Liu, Weng-Fai Wong

Abstract

Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.