Reinforcement-Learning-Based UAV Utility Maximization for Offloading Cellular Communication Systems

Base stations (BSs) with limited capacity and low coverage area face difficulties satisfying the data rate requirements of massive number of users. Unmanned aerial vehicles (UAVs) are being envisioned as a potential solution for assisting terrestrial cellular networks in Internet of Things (IoT) applications, due to their high agility, mobility, and flexibility. In this paper, we present a novel approach for maximizing a UAV's utilization to efficiently offload data traffic from terrestrial BSs. Specifically, we study the maximization of user association with the UAV by jointly optimizing UAV trajectory and user association indicators under a given data rate constraint. Since this problem is non-convex with multiple decision variables and constraints, we reformulate it into a Markov decision process and devise a reinforcement learning framework that optimizes UAV's trajectory using a low complexity state-action-reward-state-action algorithm. Our simulation results validate the analysis and provide various design insights on the optimal UAV trajectory. It is shown that the proposed design increases the average UAV-associated users by 6.75% and 14.47% compared with the Q-learning and particle swarm optimization techniques.

Paper

Full text

PDF

Reinforcement-Learning-Based UAV Utility Maximization for Offloading Cellular Communication Systems

Semantic Scholar · Computer Science · 2023

Abstract

Base stations (BSs) with limited capacity and low coverage area face difficulties satisfying the data rate requirements of massive number of users. Unmanned aerial vehicles (UAVs) are being envisioned as a potential solution for assisting terrestrial cellular networks in Internet of Things (IoT) applications, due to their high agility, mobility, and flexibility. In this paper, we present a novel approach for maximizing a UAV's utilization to efficiently offload data traffic from terrestrial BSs. Specifically, we study the maximization of user association with the UAV by jointly optimizing UAV trajectory and user association indicators under a given data rate constraint. Since this problem is non-convex with multiple decision variables and constraints, we reformulate it into a Markov decision process and devise a reinforcement learning framework that optimizes UAV's trajectory using a low complexity state-action-reward-state-action algorithm. Our simulation results validate the analysis and provide various design insights on the optimal UAV trajectory. It is shown that the proposed design increases the average UAV-associated users by 6.75% and 14.47% compared with the Q-learning and particle swarm optimization techniques.

Similar papers

© 2026 NYSGPT2525 LLC