Papers for

hardware accelerator developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Efficient method improves analog computing accuracy for AI models

IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing

Abstract: Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.

Mon 28 SeptArtificial Intelligence
The gist
Analog in-memory computing lets computers do big calculations quickly and using less energy by working in an analog way inside memory chips. But this method introduces special rounding errors because of how numbers are converted back into digital form. The authors created a new way called IMC-CLINIC that smartly adjusts how numbers are clipped and rounded to reduce these errors. Their approach works faster and more accurately than older trial-and-error methods, making AI computations more reliable.
Open → 2609.35586v1

Kv cache compression improves large language model efficiency

D-Quant: Driftable Entropy Coding for KV Cache Quantization

Abstract: The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness and ease of deployment. However, most existing methods rely on fixed-width quantization, where a $b$ bit representation is inherently limited to $2^b$ quantization levels. As the bit width decreases, the number of available levels shrinks exponentially, leading to severe information loss and rapid performance degradation. We further observe that fixed-width quantization fails to exploit the highly non-uniform distribution of KV cache. After rotation and normalization, KV values approximately follow a normal distribution, with most values concentrated near the center and only a small fraction appearing in the tails. Nevertheless, fixed-width coding allocates the same number of bits to frequent and rare symbols. Entropy coding naturally exploits such non-uniformity by assigning shorter codewords to frequent symbols and longer ones to rare symbols, substantially reducing the average number of bits required for representation. However, its variable-length output is not suited to highly parallel attention kernels, where efficient dequantization and computation rely on regular memory layouts and fixed-stride accesses. To bridge this gap, we propose \textbf{D-Quant}, a flexible KV cache quantization framework that introduces a \textbf{drift} mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.

Thu 17 SeptComputation and Language
The gist
Large language models use a key-value (KV) cache to speed up processing, but this cache takes up a lot of memory, which slows things down. Most methods to reduce this cache's size use fixed-bit quantization, which doesn't work well because it treats frequent and rare data equally and loses important information. The authors developed D-Quant, a new way to compress the KV cache that uses entropy coding to save memory without slowing down processing, by converting variable-length codes into fixed-size pieces that computers can handle efficiently.
Open → 2609.19880v1