QuantaSpike cuts energy use for large language model inference by spike-driven quantization

QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

Artificial Intelligence

Summary

Large language models are powerful but use a lot of energy because they do many heavy calculations. The authors designed a new way called QuantaSpike that uses a special kind of spiking neuron to represent information more efficiently with fewer steps. This method keeps the model’s accuracy close to normal while using much less energy. It works on several popular language models and offers a new approach for energy-efficient AI processing.

What this means in practice

  • For machine learning engineers: Run large language models in embedded or low-power devices using QuantaSpike’s spike-driven quantization to reduce energy consumption without large accuracy loss.
  • For data center operators: Cut inference energy costs for large language model services by integrating QuantaSpike to decrease power use during high-throughput deployments.

Authors

Bang Hu, Guowei Zhu, Changze Lv, Xiaoqing Zheng, Fengzhe Zhang, Fan Zhang, Wei Cao

Abstract

Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.