Papers for

high performance computing operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Partition shape strongly affects speed of distributed quantum simulations

The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations

Abstract: In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirteen torus partition geometries, and six rank densities for 39-qubit simulations on 1,024 nodes) shows that partition geometry dominates runtime. All twelve algorithms run 1.73-2.31x slower on flat partitions than on near-cubic ones despite identical data transfer, proving the slowdown stems from network delivery rather than communication volume. This penalty scales with the 3D torus partition aspect ratio (runtime $\propto a^{0.39}$, $r = 0.72$). Rank density is secondary, cutting runtime by 11% at 16 ranks per node only on compact geometries. Ultimately, requesting a near-cubic partition with 16 ranks per node roughly halves time-to-solution relative to flat partitions, which also consume 1.82x more energy. A simulator-free all-to-all microbenchmark confirms a similar geometry penalty for collective-dominated workloads.

Mon 28 SeptEmerging TechnologiesDistributed, Parallel, and Cluster ComputingPerformance
The gist
Running quantum circuit simulations on many computers depends on how the problem is split among them. The authors found that the shape of these splits affects the speed more than the amount of data shared. Using a powerful supercomputer, they showed that nearly cube-shaped partitions run about twice as fast as flat ones. This shape effect also influences how much energy the simulation uses. Their findings help optimize resource use in large quantum simulations.
Open → 2609.34369v1

HyDra improves training speed and balance for long AI sequences

HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale

Abstract: Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. Existing DCP systems either do not scale or perform poorly on mainstream models, leaving Megatron-Core (Mcore) DCP as the only option at production scale. Mcore DCP, however, sizes each degree to fit memory, which grows linearly with length while attention grows quadratically, so comparable token counts hide unequal computation. In our production 256K-context training job on more than 11K GPUs under Mcore DCP, per-rank microbatch times differ by up to 5x at similar token counts. The skew leads to a 46% pipeline bubble and a 13% data-parallel bubble. We present HyDra, a scalable load-driven DCP system. Its scheduler balances computation by pulling every rank toward one load target, and balance in turn makes that target solvable in closed form. It places sequences with lazy heaps rather than whole-pool scans. That balance asks for CP degrees larger than memory requires, so its CP engine nests an inner Ulysses group in a shallow outer ring, letting a higher degree lower computation and communication together. Evaluation at both scales shows consistent gains. On a 512-GPU testbed, HyDra raises throughput over Mcore DCP by 1.18x on average at 32K context and 2.48x at 256K. On a 2,048-GPU production job, it shrinks the pipeline bubble from 36% to 14%, cuts scheduling time by 2.3x, and raises throughput by 1.10-1.43x (avg. 1.25x) over Mcore DCP and 1.33-1.90x (avg. 1.59x) over static CP.

Mon 28 SeptDistributed, Parallel, and Cluster Computing
The gist
Training AI models on very long sequences can take uneven time because different parts handle different amounts of work. The researchers found that a common way to split this work caused big slowdowns and inefficiencies. They built HyDra, a system that smartly balances the work across all computers running the training so everything finishes more evenly. This new system made training faster and smoother, especially for very long sequences.
Open → 2609.34318v1

Managing hybrid quantum classical work improves optimization workflow control

Managing Iterative Hybrid Quantum-Classical Optimization as a First-Class Scientific Workflow

Abstract: Today's Quantum Processing Units (QPUs) are too small and too noisy to solve large combinatorial optimization problems directly, so practical hybrid solvers split a problem into pieces and iterate a decompose-solve-aggregate loop over whatever backends are available: classical heuristics, simulators, emulators, or a QPU. In practice, the loop is a driver script. It sits on top of the quantum-HPC middleware, handling task generation, provenance, recovery, and portability. Instead, we treat the loop as a scientific workflow and ask what a workflow layer adds to a generic workflow management system and QPU-sharing middleware. Two decomposition patterns from real applications, iterative consensus (ADMM) and hierarchical partitioning, turn out to stress the orchestration layer very differently: over 120 managed runs, orchestration took 78.4% of end-to-end time for the iterative pattern, almost all of it in a per-round barrier, but only 6.2% for the hierarchical one. Our workflow model adds four things a generic engine does not have: a termination predicate residing in the task graph that reads the previous round's residuals, subproblem-level recovery with warm-start and quorum-deferred aggregation, failover from a QPU to a classical replica within a round, and a provenance schema that describes CPUs, simulators and QPUs with the same fields, including the shot budget included. We report the cost of each layer on our engine, show that speculative re-execution enable runs to complete under injected failures that stall an unmanaged driver. We also use the same provenance to give per-device latency tails across simulators, emulators and IQM QPUs.

Thu 24 SeptDistributed, Parallel, and Cluster Computing
The gist
Current quantum computers are too small and noisy to directly solve big optimization problems, so they work alongside classical computers in a cycle that breaks down problems, solves pieces, and combines results. The authors looked at this cycle not just as a simple script but as a scientific workflow, adding features like how to decide when to stop, recover from errors, and switch between quantum and classical computing if needed. They tested their approach on different problem types and showed how much time the workflow management takes and how it helps handle failures. They also tracked detailed performance data from different computing devices in the same format.
Open → 2609.30532v1