Flux schedules optical switches to speed up AI model training
Flux: Optimal Scheduling of Optical Circuit Switches for LLM Training
Networking and Internet ArchitectureDistributed, Parallel, and Cluster Computing
Summary
Training large AI models needs fast and efficient data movement between computers. Optical circuit switches offer great speed but take some time to change connections, causing delays. The authors argue that planning switch changes without considering the training steps slows things down and wastes resources. They created Flux, a scheduler that plans switch changes based on the whole training work, reducing delays and buffer needs. Flux can make training up to ten times faster and greatly reduce memory use compared to older methods.
What this means in practice
- •For data center network engineers: Optimize optical circuit switch schedules to reduce AI training iteration times and buffer requirements in large-scale data centers.
- •For high performance computing teams: Improve scheduling of network resources in supercomputing clusters running AI workloads that use optical circuit switching.
Authors
Arno Troch, Seyyidahmed Lahmer, Abubakr Nada, Jeroen Famaey, Michael Peeters
Abstract
Optical Circuit Switching (OCS) offers high bandwidth density and energy efficiency for LLM training, but incurs a non-negligible reconfiguration delay. Prior work typically schedules optical circuit switches independently of compute, using aggregate traffic demand to determine which circuits to provision and when. We argue that this separation creates a fundamental inefficiency: reconfigurations that ignore the compute timeline can stall communication, resulting in low circuit utilization and large buffer requirements. In this paper, we present Flux, a scheduler that optimally schedules optical circuit switches based on the structure of the entire workload. Flux remains effective across a wide range of switching speeds by reusing circuits and amortizing reconfiguration delay behind compute and communication. We show that Flux reduces training iteration time by up to $10\times$ and peak NIC buffer requirements by more than three orders of magnitude compared to traditional periodic schedulers.