Latency aware client assignment speeds parallel split learning training
Latency-Aware Client Assignment for Parallel Split Learning With Global Sampling
Machine Learning
Summary
Training AI models using data from many different computers can be slow because some computers are slower than others. The authors found a way to decide which computer should send which part of the training data so that the slowest computers do not hold up the whole process. They developed two methods: one that finds the best assignment by solving a math problem, and another that makes quick good guesses. Their approach cuts down the total training time while still using all the data properly.
What this means in practice
- •For distributed machine learning engineers: Decrease training duration by assigning training tasks to clients considering their speeds in cross-silo split learning.
- •For cloud service schedulers: Implement latency-aware scheduling rules to minimize client-side delays when training models from distributed datasets.
Authors
Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink
Abstract
In cross-silo split learning, Parallel Split Learning with Global Sampling forms representative pooled batches when class distributions differ across clients, but ignores client delay when several clients can supply the same class. We introduce Latency Budgeted Parallel Split Learning with Global Sampling, which separates each pooled batch's integer class target from the choice of clients that supply its examples. The flow variant formulates this assignment as an integral network-flow problem and minimizes modeled client-side completion time for the current target. The fast variant uses a greedy next-completion rule to reduce schedule-construction cost. Both preserve the target stream and use every local example once per epoch. A planning rule selects between the variants while accounting for the cost of constructing both candidate schedules. On CIFAR-10, the flow variant reduces modeled training time by 6.75%, with a 0.30 percentage-point decrease in final accuracy. On Tiny ImageNet with 20 candidate classes per client, the fast variant reduces modeled time by 16.87% and reaches all four validation targets earlier than the latency-unaware baseline. Across 405 schedule comparisons, the planning rule stays within 2% of the lower realized cost in 96.54% of cases. In our evaluation, latency-aware provider assignment reduces modeled training time without changing the prescribed class targets, while the preferred variant depends on whether assignment savings outweigh schedule-construction overhead.