Papers for

software engineers for ai hardware

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

PolyCIM boosts memory chip use for faster neural network work

PolyCIM: Improving Data Reuse in Digital CIM Accelerators with Polyhedral-Based Compilation

Abstract: Digital Compute-in-Memory (CIM) presents a promising solution for accelerating deep neural networks (DNNs) through the integration of computational logic directly within memory arrays. However, mapping modern DNN operators to CIM accelerators often results in severe array underutilization, due to the strict data reuse constraints imposed by the rigid CIM array structure. We observe that data reuse in modern DNNs forms hyperplane structures often oriented along non-axial directions, rendering them invisible to conventional mapping methods that only exploit axis-aligned reuse. In this work, we propose PolyCIM, a polyhedral-based compilation framework for CIM architectures that systematically exposes and realigns these hyperplanes through affine transformations. PolyCIM provides a unified abstraction capable of efficiently representing both diverse DNN workloads and digital CIM architectures. Through data reuse exposure, computation mapping, and data movement optimization, PolyCIM generates mappings for CIM architectures that achieve superior array utilization and performance. Experimental results show that PolyCIM delivers up to $4\times$ improvement in macro utilization and $3.2\times$ speedup, effectively bridging the gap between modern DNN operators and CIM architectures.

Mon 28 SeptHardware Architecture
The gist
Computers that run tasks like recognizing images often use special chips with memory and calculators combined, called compute-in-memory (CIM) accelerators. The problem is these chips don’t work well with modern AI programs because they expect data arranged in simple, straight lines, but the real data is more complex. The authors propose PolyCIM, a smart way to reorganize data so these chips see the patterns they expect, letting them work more efficiently. This method uses math tools to transform data arrangements and speeds up processing by making better use of the chip’s resources.
Open → 2609.34351v1