Topology aware gpu load balancing speeds up moE model training
TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training
Distributed, Parallel, and Cluster ComputingMachine Learning
Summary
Training large AI models with mixtures of experts often runs into slowdowns because some GPUs get overloaded while others stay idle. The authors introduce TopoEP, a system that runs directly on GPUs and carefully balances workloads based on both the data and the hardware setup. This avoids slow communication between the computer parts and keeps all GPUs working efficiently. Their approach led to noticeable speed improvements when training on a 32-GPU cluster.
What this means in practice
- •For machine learning engineers: Improve speed and efficiency of large-scale MoE model training on GPU clusters using topology-aware load balancing.
- •For distributed systems operators: Optimize GPU resource utilization by reducing straggler effects in dynamic routing workloads across multi-node GPU setups.
Authors
Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei, Jin Fang
Abstract
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.