Thermal limits of 3D stacked memory on GPUs evaluated and mapped

Beyond HBM-on-GPU: Thermal Design Envelope for 3D Volumetric DRAM-on-GPU Integration

Emerging Technologies

Summary

GPUs used for AI and scientific computing need fast and large memory, but stacking memory chips on top of GPUs creates heat challenges. The authors studied how arranging memory vertically and adding cooling channels affects temperature and performance. They found that the height of the stack mostly controls how hot it gets, and other design choices also impact cooling efficiency. Their work defines the safe design limits for building 3D stacked memory on GPUs without overheating.

What this means in practice

  • For gpu hardware engineers: Design GPU packages with vertically stacked DRAM that balance memory capacity and thermal limits effectively.
  • For data center operators: Plan cooling solutions for GPUs with 3D stacked memory to avoid overheating and maximize training throughput.

Authors

Yukai Chen, Melina Lofrano, Khakim Akhunov, Jonas Svedas, Arjun Singh, Nathan Laubeuf, Diksha Moolchandani, Anshul Gupta, Matthew Walker, Zsolt Tokei, Geert Van der Plas, Dwaipayan Biswas, Herman Oprins, Julien Ryckaert, James Myers

Abstract

The scaling of GPUs for AI and HPC workloads is increasingly constrained by the capacity, bandwidth, and thermal limits of both 2.5D HBM-GPU and direct-stacked 3D HBM-on-GPU integration. This work establishes the thermal design envelope for 3D volumetric DRAM-on-GPU integration, in which vertically oriented DRAM dies and interleaved cooling cavities reshape heat flow and memory interfacing above the GPU. Using a package-level thermal model anchored to a consistent HBM-on-GPU baseline and driven by a realistic reticle-scale non-uniform GPU power map, we quantify the key parameters governing thermal feasibility. Stack height is the dominant limiter of peak temperature, while cooling-cavity conductivity shifts the feasible region, and mold insertion and stack orientation further modulate thermal behavior. A distributed memory-controller and network-on-chip tier introduces only a moderate thermal penalty. Although die-level parallelism increases bandwidth, the reduction in simulated training time saturates once execution becomes compute-bound. These results define a bounded co-design space across bandwidth, capacity, and thermal constraints for 3D volumetric DRAM-on-GPU integration.