Papers for

hardware accelerator designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Photonic accelerator boosts encrypted computing speed fivefold

PHAT: PHotonic Accelerator for TFHE

Abstract: Fully Homomorphic Encryption (FHE) enables secure computation on encrypted data, making it a promising solution for privacy-preserving applications in the cloud. Among various FHE schemes, FHE over the Torus (TFHE) stands out due to its support for arbitrary operations. However, its high computation and communication overhead, particularly in the Fast Fourier Transform (FFT) operations required during bootstrapping, limits its practicality for real-world applications. Conventional electronic accelerators struggle to achieve sufficient throughput due to the limitations of technology scaling and the memory-wall problem. To address these challenges, we propose PHAT, a PHotonic Accelerator for TFHE leveraging Optically-addressed Phase-Change Memory (OPCM). OPCM-based processing-in-memory systems offer high computation and communication throughput, making them well-suited for accelerating FFT operations in TFHE. However, directly mapping FFT to OPCM presents challenges such as high-precision analog computation and the high latency and energy cost of programming OPCM cells. To overcome these challenges, we introduce a novel electro-photonic accelerator architecture optimized for TFHE, featuring OPCM-based FFT units, a twiddle-stationary dataflow tailored for OPCM, and a scheduling mechanism to maximize the utilization of the FFT units. PHAT delivers $2.14\times$--$5.10\times$ speedup across four real-world TFHE workloads against the state-of-the-art ASIC accelerator. Our approach significantly enhances the performance of TFHE applications, paving the way for practical and efficient homomorphic encryption in cloud computing.

Thu 10 SeptCryptography and SecurityHardware ArchitectureEmerging Technologies
The gist
Fully Homomorphic Encryption (FHE) lets people do calculations on data without seeing the data itself, which is great for privacy. A special type called TFHE can do many types of calculations but is very slow because it needs a lot of complex math called FFT. The authors built a new photonic (light-based) device called PHAT to make these calculations much faster by using special memory technology and tailored data handling. This new approach runs TFHE tasks more than twice as fast as the best existing hardware, making encrypted cloud computing more practical.
Open 2609.11613v1

Quantization moves perform differently depending on model state and order

Contextual Utility of Quantization Moves in Extreme Low-Bit LLMs

Abstract: Post-training quantizers select finite code changes using reconstruction proxies or local loss approximations, but the utility of a quantization move depends on the state through which it is executed. We identify two sources of this contextual dependence. First, the displacement of the move matters: evaluating the gradient at the move midpoint captures curvature accumulated along the move that a current-state linearization omits. Across frozen two-bit moves from Llama-3.2 models, midpoint evaluation predicts the direction of exact endpoint loss changes substantially more accurately than current-state gradients. Second, moves interact: exhaustive lattices of legal quantized states are well approximated by quadratic pseudo-Boolean functions, yet their small pairwise components can determine Pareto fronts and cause different evaluation functionals to prefer opposite directions. These effects explain failures of reconstruction-optimal code re-selection and additive composition. Reading each move at its own midpoint repairs the local selection step and improves downstream accuracy and held-out perplexity, while larger supports require evaluating exact endpoints from the state actually reached. Exact-endpoint beam search finds sparse changes that dominate much larger one-shot updates, and repricing the same moves after intervening changes produces widespread sign reversals. These results show that quantization utility is contextual at the granularity of a few moves: reliable construction must evaluate finite changes along their own paths and compose them from the evolving quantized state.

Wed 9 SeptDatabases
The gist
When compressing huge language models to use fewer bits, the impact of each small change depends on the current state of the model and the order in which changes happen. The authors show that measuring the effect of a change at its midpoint better predicts how it will affect model quality than just looking at the starting point. They also found that moves interact in complex ways, meaning that checking changes in isolation can be misleading. Their methods improve the accuracy of tiny-bit models by carefully evaluating these changes along their actual paths.
Open 2609.09867v1

Transformer softmax calculation sped up with new low-bit quantization method

EFQ-Softmax: Exp-Free Quantization for Softmax

Abstract: Low-bit attention accelerates Transformer inference by moving the $QK^\top$ and $PV$ matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates shifted-score exponentials in higher precision, forms a temporary probability block, and quantizes it before low-bit $PV$ multiplication. This exp-then-quantize path creates a mismatch between a high-precision probability producer and a low-bit matrix consumer. We propose EFQ-Softmax (Exp-Free Quantization for Softmax), a low-bit probability-generation method that directly maps shifted attention scores to block-scaled E2M1 operands. For each microscaling block, EFQ-Softmax selects an exponent-only scale from the local maximum, maps the shifted scores to a normalized residual domain, and generates nonnegative E2M1 probability codes using a single affine rule. The resulting operand is used consistently in both the $\widetilde{P}V$ numerator update and the $\widetilde{P}\mathbf{1}$ denominator update. The FlashAttention-style row-maximum update, historical rescaling, high-precision accumulation, and final normalization remain unchanged. We evaluate end-to-end quality on Qwen3-8B, Qwen3-VL-8B-Instruct, and WAN2.2-TI2V-5B, and separately measure kernel-level performance on the A5 vector unit. EFQ-Softmax improves the Qwen3-8B seven-task mean from 0.6749 with MXFP4 to 0.6773 and the Qwen3-VL nine-task mean from 0.7826 to 0.8000. On WAN2.2, it maintains temporal consistency and visual quality comparable to the FP16 and MXFP4 baselines under VBench. On the A5 vector unit, EFQ-Softmax reduces the vector-stage latency of the fused probability-generation kernel by 40.33% on average across sequence lengths from 16K to 128K. These results show that direct low-bit probability generation can replace the conventional exp-then-quantize path while preserving end-to-end model quality.

Wed 9 SeptMachine Learning
The gist
Transformer models use a step called softmax that usually needs higher precision math, which slows things down. The authors introduce EFQ-Softmax, which skips the complex exponential step and instead works directly with simpler low-bit numbers. This approach matches the precision across calculations better and makes the process faster without losing quality. They tested it on several language and vision models, finding small improvements in accuracy and a significant speed boost on specialized hardware.
Open 2609.09721v1

Improving floating point quantization noise prediction in matrix multiplication

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

Abstract: We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $κ$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $κ^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.

Tue 8 SeptMachine LearningArtificial Intelligence
The gist
When computers do many calculations with limited precision numbers, they introduce tiny errors called quantization noise. The authors created a new mathematical way to understand how this noise behaves in matrix multiplication when using different types of number formats. They derived a formula that predicts the noise's effect and found the best possible limit on how much noise could be reduced through certain transformations. Using this insight, they developed a new technique called KBBQ that better controls this noise in floating-point 4-bit formats, leading to more accurate results without slowing down the computation.
Open 2609.08135v1