Consensus federated learning matches centralized training for robot vision models
Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models
Robotics
Summary
Training robots to understand images, language, and actions often needs lots of data collected in many places, which is hard to gather in one spot. The authors developed Co-VLA, a method that lets many robots train a shared brain without sharing their private data by reaching agreement through a consensus process. This approach works well whether training the entire model or fine-tuning parts of it, performing as good as if all data were gathered centrally. Their method helps overcome challenges of different robots having different types of data.
What this means in practice
- •For robotics engineers: Train robot vision-language-action models collaboratively across multiple robots without sharing raw data.
- •For machine learning teams: Perform decentralized training or fine-tuning on distributed data with heterogeneous distributions using consensus-based federated learning.
Authors
Haolong Li, Guner Dilsad Er, Michael Muehlebach, Joerg Stueckler
Abstract
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings.