Partition shape strongly affects speed of distributed quantum simulations
The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations
Emerging TechnologiesDistributed, Parallel, and Cluster ComputingPerformance
Summary
Running quantum circuit simulations on many computers depends on how the problem is split among them. The authors found that the shape of these splits affects the speed more than the amount of data shared. Using a powerful supercomputer, they showed that nearly cube-shaped partitions run about twice as fast as flat ones. This shape effect also influences how much energy the simulation uses. Their findings help optimize resource use in large quantum simulations.
What this means in practice
- •For quantum computing engineers: Optimize cluster partition shapes to speed up quantum circuit simulations and reduce energy in large-scale distributed environments.
- •For high performance computing operators: Improve network usage in multi-node simulations by choosing near-cubic partitioning, lowering runtime and energy costs on torus networks.
Authors
Yikai Mao, Yuan He, Shaowen Li, Masaaki Kondo
Abstract
In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73-2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime $\propto a^{0.39}$, $r = 0.72$). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free all-to-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads.