AI-RAN systems share GPU power to speed up foundation model training
Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training
Distributed, Parallel, and Cluster ComputingNetworking and Internet Architecture
Summary
Cell sites with special GPUs often have unused computing power. The authors studied how this extra power can be safely used to train big AI models without slowing down the main cell tasks. They designed a system called Weaver that manages GPU use cleverly to fit AI training around cell operations. Weaver lets multiple cell sites share their spare power, increasing how much AI work gets done. Tests showed this method greatly improved training speed without hurting cell service.
What this means in practice
- •For cell network operators: Run foundation model training jobs on idle GPU resources in AI-enabled cell sites without affecting network performance.
- •For edge ai developers: Increase large model training throughput by utilizing spatially bursty GPU resources spread across multiple cell sites.
Authors
Leyang Xue, Tianxin Wang, Xin Zhe Khooi, Jiaxun Yang, Dheeraj Mahendiran, Yufeng Xia, Mun Choon Chan, Myungjin Lee, Mahesh K. Marina
Abstract
The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. We explore using this spare capacity for decentralized training of foundation models (FMs), one of the most compute-intensive AI workloads. We present the first characterization of spare GPU capacity in AI-RAN systems at both micro-scale--across transmission slots within a cell site--and macro-scale--across sites. Our analysis finds that 40-85% of GPU capacity is unused; although this capacity is temporally bursty at individual sites, it is spatially complementary across sites. To safely and efficiently harness these resources, we present Weaver, a system that opportunistically trains FMs alongside latency-critical RAN workloads without degrading RAN performance. Weaver adopts a RAN-first design: a spare-compute controller integrated into the MAC scheduler uses compute-aware scheduling to smooth RAN GPU demand and exposes more usable spare GPU capacity. A two-level elastic training framework then adapts to dynamic, heterogeneous spare capacity within and across sites. Experiments on an O-RAN-aligned system prototype show that Weaver creates up to 4.9x more usable spare compute and utilizes up to 83% of the available spare capacity. On a multi-site testbed, Weaver improves training throughput by 2.1-3.7x over baseline approaches.