Adaptive client clustering improves federated learning in edge networks
Adaptive Client Clustering and Coordination for Federated Learning Workflow Management in Edge Networks
Distributed, Parallel, and Cluster Computing
Summary
Federated learning lets many devices work together to train a machine learning model without sharing their data directly, but this gets tricky when devices differ a lot and network delays vary. The authors propose A-CoDa, a method that groups similar devices into clusters and carefully decides which ones should participate in learning steps to speed up the process and keep the model accurate. They also design a scheduler to manage tasks efficiently, respecting the order needed for learning to progress. Tests on various real-world problems show that their approach cuts down training time while maintaining good results.
What this means in practice
- •For edge network operators: Manage device groups and schedule federated learning tasks to reduce training time for connected sensors and devices with varying data and network conditions.
- •For mobile application developers: Use adaptive client selection to maintain model accuracy and timely updates when training models across diverse smart devices with uneven data distributions and network delays.
Authors
Jieping Luo, Qiyue Li, Yuxuan Chen, Hang Qi, Jiaying Yin, Jingjin Wu, Qian Wang
Abstract
Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinate heterogeneous clients, non-IID data, fluctuating communication latency, and precedence-constrained tasks under service-level completion requirements. These coupled factors make participant management central to both time-totarget performance and learning stability. This paper proposes A-CoDa, an adaptive clustered coordination framework for managing dependent FL flows. A-CoDa first uses label-distribution divergence (LDD)-based greedy-balanced clustering to construct statistically coherent and size-aware client groups, which serve as a scalable management abstraction. Building on this structure, we design FedMIX, an uncertainty-aware intra-/inter-cluster participation mechanism that ranks clients by a loss-latency-uncertainty utility and adaptively controls cross-cluster probing according to training progress and latency conditions. A dependency-aware DAG scheduler then orchestrates layer-wise task execution so that parallelism and precedence constraints are jointly respected. We further provide a convergence analysis that frames the result as a sufficient loss-domain design bound, explicitly relating the attainable error floor and sufficient communication rounds to LDDinduced sampling mismatch, residual distribution shift, local-SGD drift, stochastic variance, and adaptive probing budgets. Experiments on handwriting, wearable-sensing, product-image, and medical-imaging tasks evaluate A-CoDa under dependent FL workflows and demonstrate its effectiveness in reducing end-toend completion time while maintaining competitive accuracy.