LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network Slicing

2026-08-03Networking and Internet Architecture

Networking and Internet Architecture
AI summary

The authors study how to manage resources in satellite networks that orbit close to Earth and connect with 5G and future 6G systems. They create a smart scheduler using deep reinforcement learning that works on two timescales: a slower one to pick overall policies and a faster one to handle quick user requests. Their method helps keep delays low for important mission-critical traffic without starving other users. Simulations show it balances network capacity and delay better than traditional approaches, making it useful for future satellite-based mobile networks.

Low Earth Orbit (LEO)Non-Terrestrial Networks (NTN)5G6GOpen Radio Access Network (O-RAN)deep reinforcement learning (DRL)resource managementqueuing delayMarkov Decision ProcessTD3 agent
Authors
Víctor Vilchez, Tiago P. C. de Andrade, Edward Hinojosa, Edmundo Madeira, and Carlos A. Astudillo
Abstract
The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base station mobility, and channel non-stationarity, complicating radio resource management of heterogeneous network slices. In this paper, we propose a deep reinforcement learning (DRL) meta-scheduler for twin-timescale resource allocation. Our solution adopts a decoupled Open Radio Access Network (RAN) architecture, in which a strategic 100 ms meta-scheduler selects scheduling policies for the different network slices using stale telemetry, while a fast-timescale MAC packet scheduler processes per-TTI user requests. The resulting Markov Decision Process captures non-stationary orbital dynamics and heterogeneous SLAs constraints via a TD3 agent. Simulation results under varying traffic load show that, unlike other solutions, the proposed meta-scheduler explicitly trades a statistically insignificant 1% capacity fraction (p > 0.05) to strictly bound the variance and overall magnitude of RLC-layer queuing delay for Mission-Critical (MC) traffic. Crucially, it enforces this isolation without inducing the broadband slice starvation characteristic of standard maximum-CQI heuristics, establishing a robust foundation for 6G O-RAN NTN resource allocation.