Papers for

gpu cluster operators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

VarioPath speeds up GPU cluster data sharing by managing PCIe traffic

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

Abstract: AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.

Mon 28 SeptDistributed, Parallel, and Cluster ComputingNetworking and Internet Architecture
The gist
Sharing information quickly between many GPUs is important for running large AI language models smoothly. On systems using PCIe connections, the data transfer can get jammed because many transfers compete for the same routes. The authors designed VarioPath, a method that smartly plans these data transfers based on the computer setup and current demands. This approach helps reduce delays and speeds up tasks like AI model inference on GPU clusters.
Open → 2609.34340v1

Charm++ runtime enables fast no-restart resizing for cloud HPC

No-Restart Elasticity in an Adaptive Runtime System for Cloud-Native HPC

Abstract: Exploiting discounted spot instances for HPC requires an application to change its resource allocation at runtime, shrinking ahead of an interruption and expanding onto replacement capacity. Existing elasticity mechanisms implement rescaling as a full process teardown followed by a cold restart at the new processor count, and the restart stage accounts for up to 95% of the total overhead, growing with node count and increasing substantially on GPUs due to CUDA context initialization. In this paper, we present a no-restart rescaling mechanism for the Charm++ runtime system in which surviving processes never exit: a rescaling operation consists of a membership negotiation with a lightweight external coordinator, reconciliation of the UCX communication endpoints with the new cluster view, and a return to the top of the runtime initialization path that preserves live application state. This reduces the cost of a rescaling operation, excluding the load balancing step that any rescaling model requires, from multiple seconds to 8--15ms on CPUs and 7--11ms on GPUs, at 4 to 32 instances. Because processes survive, GPU device state persists in place, eliminating the checkpointing daemons previously required for GPU elasticity, and a launcher-independent bootstrap mechanism removes the dependence on supervised process managers, which are incompatible with spot instance interruptions. Integrated with an existing spot instance management framework, the mechanism cuts the end-to-end overhead of eight simultaneous interruptions to 0.2% of runtime on CPUs and 0.6% on GPUs, and raises the rate at which a job can be rescaled below 1% overhead by a factor of seven.

Sat 26 SeptDistributed, Parallel, and Cluster Computing
The gist
Running big computing jobs in the cloud often means using cheap but unreliable machines that can disappear suddenly. The authors present a way for these jobs to shrink or grow without the usual long pause to restart everything. Their method keeps all running processes alive and just updates connections and states quickly, saving seconds down to milliseconds. This helps programs using GPUs too, avoiding slow restarts related to GPU setup. Overall, their approach cuts downtime drastically, making cloud computing with spot machines smoother and cheaper.
Open → 2609.32217v1

Compass-abs reduces gpu cluster resource fragmentation for faster deep learning

COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

Abstract: With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quantifying fragmentation and developing scheduling algorithms that alleviate its impact. However, existing fragmentation measures break down in the absence of workload distribution information, while current schedulers cannot continuously maintain resource fragmentation at a low level. To tackle these problems, we first introduce Scheduler-Induced Fragmentation (SIF), a metric built on the notion of partial-nodes that is independent of historical workload knowledge. We then propose COMPASS-ABS, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS), whose construction fully leverages the topological alignment between dominant workload size and node capacity. Moreover. We also prove that it ensures SIF is bounded by $\frac{2}{N}$ under a workload composition condition that matches both theory and production. Evaluations implemented on a physical cluster and a simulated cluster demonstrate COMPASS-ABS effectiveness at improving resource utilization, reducing DLT job completion time by reducing fragmentation.

Wed 16 SeptDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Large clusters of GPUs are often used to train deep learning models, but their resources can become scattered and inefficiently used, causing delays. The authors introduce a new way to measure this problem, called Scheduler-Induced Fragmentation, that doesn’t rely on knowing past workloads. They then propose a method named COMPASS-ABS to keep the GPU resources organized by aligning job sizes with node capacities. Their approach helps make better use of GPUs and shortens the time it takes to complete training jobs.
Open → 2609.18519v1