CoeF-SFL cuts communication in split federated learning without losing accuracy
CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency
Artificial Intelligence
Summary
Training AI models together while keeping data private is tricky because it often requires a lot of back-and-forth communication. The paper shows that previous methods which tried to reduce communication made the client and server learn different goals, leading to worse teamwork. The authors propose a new way called CoeF-SFL that sends updates less often but adjusts for stale information so the learning stays coordinated. Their method works well across different tasks like vision and language and beats older approaches while using less communication.
What this means in practice
- •For mobile app developers: Enable collaborative model training on resource-limited devices by reducing communication needs without losing model accuracy.
- •For edge computing engineers: Design distributed AI systems that efficiently share computation and model updates between edge devices and servers.
Authors
Junwoo Bae, Jin-Hyun Ahn
Abstract
Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at https://anonymous.4open.science/r/CoeF-SFL-2686/README.md