Topology aware gpu load balancing speeds up moE model training

TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

Distributed, Parallel, and Cluster ComputingMachine Learning

Summary

Training large AI models with mixtures of experts often runs into slowdowns because some GPUs get overloaded while others stay idle. The authors introduce TopoEP, a system that runs directly on GPUs and carefully balances workloads based on both the data and the hardware setup. This avoids slow communication between the computer parts and keeps all GPUs working efficiently. Their approach led to noticeable speed improvements when training on a 32-GPU cluster.

What this means in practice

Authors

Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei, Jin Fang

Abstract

Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.