Federated adaptive knowledge distillation cuts communication for large language models

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

Artificial Intelligence

Summary

Large language models are powerful but usually need to be trained using lots of shared data, which raises privacy issues. The authors work on federated learning, which lets many devices train a shared model without sharing private data. They propose a new method called FLoKD that sends only the most important model information during training to reduce communication over slow wireless networks. Their method sends compressed signals instead of full parameters or detailed outputs, and smartly picks parts of data and model to share. This approach significantly lowers communication needs while keeping good performance on language tasks.

What this means in practice

  • For mobile network engineers: Optimize federated model updates over limited bandwidth wireless links by selectively transmitting important model components to reduce communication costs.
  • For cloud service operators: Enable privacy-preserving collaborative language model fine-tuning with lower communication overhead by employing adaptive knowledge distillation techniques.

Authors

Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi

Abstract

Large language models (LLMs) have demonstrated strong capabilities across a wide range of natural language processing tasks. However, conventional fine-tuning typically relies on centralized data collection, bringing in privacy concerns. Federated learning (FL) enables collaborative LLM fine-tuning without sharing raw client data, but its deployment over bandwidth-constrained wireless networks is hindered by the communication overhead of model-parameter transmission. Although Low-Rank Adaptation (LoRA) reduces the number of trainable parameters, its communication cost still increases with model scale. Knowledge distillation avoids parameter sharing via output logits, but token-level logits in LLMs incur high communication cost due to sequence length and vocabulary size. Reducing logits lowers the cost but weakens supervision and degrades accuracy. To address these limitations, we propose FLoKD, an adaptive knowledge-distillation framework for federated LoRA fine-tuning of LLMs over wireless networks, which communicates intermediate LoRA activations as the distillation signal rather than logits or full parameters. Since transmitting all blocks over the entire public dataset remains costly, we further propose a transformer block importance scoring framework that selectively transmits the most informative blocks, and two dataset selection strategies that discard public samples deviating from the local data distribution and prioritise those most informative for distillation. Extensive experiments across multiple generative language datasets, including WikiText-103, PTB, and Dialog, demonstrate that our proposed framework reduces communication overhead by 50-65% while achieving rapid convergence to competitive perplexity compared to baselines.