Charm++ runtime enables fast no-restart resizing for cloud HPC
No-Restart Elasticity in an Adaptive Runtime System for Cloud-Native HPC
Distributed, Parallel, and Cluster Computing
Summary
Running big computing jobs in the cloud often means using cheap but unreliable machines that can disappear suddenly. The authors present a way for these jobs to shrink or grow without the usual long pause to restart everything. Their method keeps all running processes alive and just updates connections and states quickly, saving seconds down to milliseconds. This helps programs using GPUs too, avoiding slow restarts related to GPU setup. Overall, their approach cuts downtime drastically, making cloud computing with spot machines smoother and cheaper.
What this means in practice
- •For cloud infrastructure engineers: Implement adaptive resource scaling in HPC jobs to minimize downtime and cost using spot instances without restarting processes.
- •For gpu cluster operators: Maintain GPU device state during cluster size changes to reduce overhead and avoid checkpointing when reallocating resources.
- •For financial modeling teams: Speed up cloud-based simulations by rescaling computing resources instantly, reducing interruption time during spot instance changes.$Commercial implications: Enables cloud simulation services to offer faster and cheaper elastic compute with minimal downtime for financial firms.
Authors
Aditya Bhosale, Laxmikant Kale
Abstract
Exploiting discounted spot instances for HPC requires an application to change its resource allocation at runtime, shrinking ahead of an interruption and expanding onto replacement capacity. Existing elasticity mechanisms implement rescaling as a full process teardown followed by a cold restart at the new processor count, and the restart stage accounts for up to 95% of the total overhead, growing with node count and increasing substantially on GPUs due to CUDA context initialization. In this paper, we present a no-restart rescaling mechanism for the Charm++ runtime system in which surviving processes never exit: a rescaling operation consists of a membership negotiation with a lightweight external coordinator, reconciliation of the UCX communication endpoints with the new cluster view, and a return to the top of the runtime initialization path that preserves live application state. This reduces the cost of a rescaling operation, excluding the load balancing step that any rescaling model requires, from multiple seconds to 8--15ms on CPUs and 7--11ms on GPUs, at 4 to 32 instances. Because processes survive, GPU device state persists in place, eliminating the checkpointing daemons previously required for GPU elasticity, and a launcher-independent bootstrap mechanism removes the dependence on supervised process managers, which are incompatible with spot instance interruptions. Integrated with an existing spot instance management framework, the mechanism cuts the end-to-end overhead of eight simultaneous interruptions to 0.2% of runtime on CPUs and 0.6% on GPUs, and raises the rate at which a job can be rescaled below 1% overhead by a factor of seven.