Papers for
hardware engineers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Analogue memory hardware powers bio-inspired probabilistic decision making
Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1
Abstract: Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could they be replicated in computing systems? This abstract discusses a biologically grounded framework in which noisy neural and synaptic dynamics perform inference and learning via stochastic sampling from an internal energy function, capturing uncertainty over latent states and model parameters through neural and synaptic variability, respectively. This enables approaches such as predictive coding networks to account for epistemic uncertainty via Markov chain Monte Carlo sampling. Drawing a parallel between intrinsic noise in biological systems and electrical noise in emerging probabilistic analogue memory technologies, we highlight how analogue in-memory computing hardware naturally emerges as the solution for massively scalable and energy-efficient probabilistic inference.
Controller design improves long-span error correction for high-bandwidth memory
REACH: Controller-Managed Long-Span ECC for HBM AI Inference
Abstract: High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a comparable code rate, but a direct implementation couples small accesses to span-wide state and requires costly decoding at HBM bandwidth. Read-dominated LLM decode offers a favorable setting: sequential reads support span aggregation, while sparse writes limit parity-update traffic. This paper presents REACH, a controller microarchitecture that uses established inner codes to correct common errors and identify unresolved chunks, reserving a long outer code for known-erasure repair. Differential parity bounds write traffic, and a co-designed endpoint preserves 32\,B transactions without an extra data burst. Ramulator2 sustains 1.88\,TB/s of application traffic at the highest error stress, while separate full-interface sizing supports a 2.69\,TB/s application target using ASAP7-synthesized kernels. At this analytical target, REACH's nominal composition uses 55.8\% less controller area and 57.7\% less modeled power than the evaluated mean-work direct-long design, showing the benefit of reserving long-span recovery for exceptional requests.
Symboliclight v2 cuts energy use for language models with new hardware design
SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference
Abstract: SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement the 194M-parameter model on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution. Across three same-checkpoint FPGA implementations at 175 MHz, active-row weight gathering and valid-state KV loading raise decode throughput from 474.6 to 643.2 tokens/s for a 32-token prefix and 128 outputs. Estimated gross card energy falls from 0.06087 to 0.04407 J per generated token, a 27.6% reduction. Complete-request energy, including prefill, falls by 24.4-27.7% across three prefix lengths. An independent idle split attributes 82.8% of gross card energy to loaded idle, explaining the benefit of shorter token latency. Against the recorded RTX 5090 compiled-FP32 baseline, integer FPGA execution uses 89.1% less estimated card energy during short-context decode; arithmetic precisions differ, and the GPU baseline is not the lowest-energy tested configuration. On four Cortex-A76 cores of a ROCK 5T, complete requests reach 65.4 tokens/s at 9.80 W and 0.151 J per generated token at the adapter's AC input. These results connect event sparsity to omitted computation and data movement. The mechanisms also support other dedicated V2 implementations: increasing throughput by a greater factor than active power lowers energy per generated token. Evaluation holds the deployed checkpoint fixed; its quality trails a same-budget dense control, so the results do not establish equal-quality efficiency.
Two-transistor one-rram chip enables more reliable in-memory computing
OTTER - Two Transistor - One RRAM Architecture for Reliable In-Memory-Computing in 28 nm CMOS Technology
Abstract: This work presents OTTER, a 28 nm CMOS platform co-integrated with TaOx-based valence-change mechanism (VCM) RRAM, demonstrating a two-transistor-one-memristive-device (2T1R) architecture for reliable in-memory computing. The 2T1R cell combines a low-drive-current (LD) transistor and a high-drive-current (HD) transistor in parallel, providing dedicated bias paths for SET programming and RESET operation, respectively. Through systematic experimental and simulated comparison of various transistor-pairing configurations using the physical compact model JART VCM Rth, design guidelines for transistor sizing are derived, establishing the minimum RESET transistor W/L required for complete RESET as a function of the SET current compliance. The 2T1R cell is further characterized under pulse-based programming, demonstrating multilevel analog conductance tuning with narrow, well separated conductance states across six programmable levels. An analog content-addressable memory (aCAM) design based on the same 2T1R cell is additionally analyzed at the circuit level, evaluating trade-offs between top- and bottom-connected RRAM comparator configurations. A hardware implementation of compute-in-memory (CIM) multiply-and-accumulate (MAC) operations is further demonstrated on a 15 x 15 2T1R crossbar array.
Digital chip with 27,648 spins solves complex problems fast
A 28nm 27,648-Spin Multichip Digital Ising Accelerator with Pegasus Connectivity
Abstract: We present a 28nm digital Ising accelerator with 27,648 spins across four chips. A time-multiplexed spin-update array with local SRAM and scheduled interchip transfers delivers 41.5G updates/s at 1.2pJ/update. Degree-15 Pegasus connectivity and 10b coefficients increase native connectivity and precision over degree-8, 5b multichip annealers. Programmable couplings support optimization and probabilistic logic, with a four-chip planted-MaxCut trace reaching the solution for 27,069 nodes in 3.3$μ$s.
Photonic reservoir computing overcomes readout size limits with compression
Photonic reservoir computing with dimensionally compressed readout
Abstract: This work addresses a hardware constraint in reservoir computing: the limited size of the readout layer imposed by systems with a physical readout. We investigate a strategy to accommodate this constraint based on random projection, which compresses high-dimensional reservoir states into a lower-dimensional subspace while preserving key properties of the source space and information- processing capabilities. To evaluate this approach, we compare a small, standalone time delay reservoir against a larger configuration whose output is projected down to match the same restricted readout dimension. Using task-independent metrics, we demonstrate that the distribution of information-processing capacities may differ between the two configurations, even at identical readout sizes. Furthermore, we perform a comprehensive hyperparameter scan to assess how both systems behave under varying physical regimes. Finally, we benchmark this approach on the standard NARMA10 task, showing that the random projection framework can yield superior performance compared to a standalone constrained reservoir, within specific compression range. These results provide a scalable pathway to bypass physical readout bottlenecks in hardware-based reservoir computing.
Analog compute in memory improves accuracy despite noise
NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-Memory
Abstract: Analog compute-in-memory (CIM) enables energy-efficient model acceleration, but its reliance on ADC-based readout, which directly quantizes noisy column currents, makes inference accuracy highly sensitive to analog read noise, active-row scaling, and ADC precision. In this paper, we present NOVA-CIM, a noise- and correlation-tolerant stochastic interface for analog CIM by replacing multi-bit ADC readout with random-reference 1-bit sensing and reconstructing results through lightweight counting. By converting column currents into comparison probabilities, this probability-domain readout averages zero-mean dynamic read noise over stochastic samples while reducing dependence on high-resolution ADCs. We provide a unified robustness analysis showing that dynamic read noise is suppressed through temporal averaging and that spatial input-bitstream correlation increases instantaneous current variance rather than introducing first-order MAC bias. MAC-level experiments and end-to-end evaluation on ViT-Base validate the analysis: under read noise, Top-1 accuracy remains 84.48% near the 84.51% bfloat16 (BF16) baseline; under stochastic number generator (SNG) reuse, MAC bias stays near zero while root-mean-square error (RMSE) and stochastic cross-correlation (SCC) grow as predicted.