Cascade Reinforcement Learning with State Space Factorization for O-RAN-based Traffic Steering

We study the Traffic Steering (TS) problem in Open Radio Access Network (O-RAN), leveraging its RAN Intelligent Controller (RIC), in which RAN configuration parameters of cells can be jointly and dynamically optimized in near-real-time. To address the TS problem, we propose a novel Cascade Reinforcement Learning (CaRL) framework, where we propose state space factorization and policy decomposition to mitigate the need for large complex models and well-labeled datasets. For each sub-state space, an RL sub-policy is trained to optimize the Quality of Service (QoS). To apply CaRL to new network areas, we propose a knowledge transfer approach to initialize a new sub-policy based on knowledge learned by the trained policies. To evaluate CaRL, we build a data-driven and scalable RIC Digital Twin (DT) that is modeled using real-world data, including network setup, user geo-distribution, and traffic demand, among others, from a tier-1 RAN operator. We evaluated CaRL in two DT scenarios representing two different US cities and compared its performance with business-as-usual policy as a baseline and other competing optimization approaches (i.e., heuristic and Q-table algorithms). Furthermore, we have conducted a field trial with the RAN operator to evaluate the performance of CaRL in two areas in the Northeast US regions.

Paper

Similar papers

© 2026 NYSGPT2525 LLC