Energy-Latency Trade-offs in O-RAN with Distributed Baseband Processing and AI Inference
2026-08-03 • Networking and Internet Architecture
Networking and Internet Architecture
AI summaryⓘ
The authors study how to best split and place tasks in Open Radio Access Networks (O-RAN) to balance energy use and delay. They build a detailed model that measures both energy consumption and latency, including costs from using AI. Using this model, they create an optimization problem to decide where to put processing and AI tasks under different conditions like network load and energy limits. Their results show how user needs and network stress affect the best setup, revealing trade-offs between saving energy and keeping delays low. This work helps guide how to deploy O-RAN systems efficiently while supporting AI-driven services.
Open Radio Access NetworkO-RANfunctional splitslatencyenergy efficiencybaseband processingAI inferenceenergy consumption modeloptimizationnetwork load
Authors
Urooj Tariq, Rishu Raj, Shashi Raj Pandey, Merim Dzaferagic, Petar Popovski, Dan Kilper
Abstract
The Open Radio Access Network (O-RAN) architecture introduces flexible functional splits and open interfaces that enable distributed and centralized deployment of baseband processing. While this flexibility offers opportunities for improved resource utilization, it also introduces fundamental trade-offs between energy efficiency and latency. In this paper, we develop a throughput-based end-to-end energy consumption model for O-RAN and extend it by incorporating detailed latency modeling and application-specific Artificial Intelligence/Machine Learning inference costs. The proposed end-to-end modeling framework provides a general representation of processing, transport, and inference-related energy and delay across the access, metro, and long-haul network segments. Building on this general model, we formulate an optimization problem that selects the placement of baseband processing and AI inference tasks across candidate O-RAN configurations to analyze energy-latency tradeoffs under network load, server frequency, and energy-budget constraints. Using representative hardware platforms and realistic traffic assumptions, we evaluate multiple baseband processing placements corresponding to different O-RAN functional configurations. Our results reveal how user quality of service requirements and network load conditions jointly determine the optimal placement of baseband processing and AI inference tasks, highlighting the inherent trade-off between energy efficiency and latency. The analysis provides practical insights for latency-aware and energy-efficient O-RAN deployments supporting emerging AI-driven services.