VarioPath speeds up GPU cluster data sharing by managing PCIe traffic

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

Distributed, Parallel, and Cluster ComputingNetworking and Internet Architecture

Summary

Sharing information quickly between many GPUs is important for running large AI language models smoothly. On systems using PCIe connections, the data transfer can get jammed because many transfers compete for the same routes. The authors designed VarioPath, a method that smartly plans these data transfers based on the computer setup and current demands. This approach helps reduce delays and speeds up tasks like AI model inference on GPU clusters.

What this means in practice

  • For gpu cluster operators: Improve large AI model inference speeds by optimizing data transfers on PCIe GPU clusters with VarioPath's scheduling method.
  • For data center network engineers: Design and configure GPU system data paths to reduce PCIe link congestion during heavy all-to-all communications using topology- and demand-aware schedules.

Authors

Yao Fei, Jin Fang, Size Zheng, Gongming Zhao, Hongli Xu, Jiacheng Zhu, Zhijing Xin

Abstract

AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.