Offloading gpu communication tasks frees cores for faster ai training
HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training
Networking and Internet Architecture
Summary
Training large AI models takes a lot of work on GPUs, which have special parts called SMs to do the hard math. Normally, these SMs also help move data between GPUs during training, which slows down the math work. The authors found a way to let a different part of the GPU, called the DMA engine, handle data movement instead, freeing up SMs to focus only on math. Their system, HOCCL, manages this approach and keeps communication fast while speeding up overall training by up to 5%.
What this means in practice
- •For machine learning engineers: Increase AI model training speed by reducing GPU core involvement in internal communications.
- •For gpu system architects: Design GPU communication frameworks that offload data movement to DMA engines, boosting computation efficiency.
Authors
Yao Fei, Gongming Zhao, Hongli Xu, Jin Fang, Jiacheng Zhu, Shuo Xu, Kun Huang, Zhuolong Yu
Abstract
Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.