Compass-abs reduces gpu cluster resource fragmentation for faster deep learning
COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads
Distributed, Parallel, and Cluster ComputingMachine Learning
Summary
Large clusters of GPUs are often used to train deep learning models, but their resources can become scattered and inefficiently used, causing delays. The authors introduce a new way to measure this problem, called Scheduler-Induced Fragmentation, that doesn’t rely on knowing past workloads. They then propose a method named COMPASS-ABS to keep the GPU resources organized by aligning job sizes with node capacities. Their approach helps make better use of GPUs and shortens the time it takes to complete training jobs.
What this means in practice
- •For gpu cluster operators: Manage shared GPU clusters more efficiently to reduce wasted resources and speed up deep learning job completion.
- •For cloud service providers: Improve scheduling to increase GPU cluster utilization, helping offer faster AI training services to customers.$Commercial implications: Enables cloud providers to offer enhanced GPU cluster rental services by reducing inefficiencies in resource use.
Authors
Yukai Zhou, Hongfan Wu
Abstract
With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quantifying fragmentation and developing scheduling algorithms that alleviate its impact. However, existing fragmentation measures break down in the absence of workload distribution information, while current schedulers cannot continuously maintain resource fragmentation at a low level. To tackle these problems, we first introduce Scheduler-Induced Fragmentation (SIF), a metric built on the notion of partial-nodes that is independent of historical workload knowledge. We then propose COMPASS-ABS, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS), whose construction fully leverages the topological alignment between dominant workload size and node capacity. Moreover. We also prove that it ensures SIF is bounded by $\frac{2}{N}$ under a workload composition condition that matches both theory and production. Evaluations implemented on a physical cluster and a simulated cluster demonstrate COMPASS-ABS effectiveness at improving resource utilization, reducing DLT job completion time by reducing fragmentation.