Analysis of Reinforcement Learning Schemes for Trajectory Optimization of an Aerial Radio Unit
This paper introduces the deployment of unmanned aerial vehicles (UAVs) as lightweight wireless access points that leverage the fixed infrastructure in the context of the emerging open radio access network (O-RAN). More precisely, we introduce the aerial radio unit (A-RU) that dynamically serves an underserved area and connects to the distributed unit (ODU) via a wireless fronthaul between the UAV and the closest fixed network infrastructure tower. In this paper we employ artificial intelligence (AI) for determining the UAV trajectory for serving User Equipment (UEs) while maintaining the fronthaul connectivity to the O-DU at the same time in a multiple-input multiple-output (MIMO) fading channel. We first formulate the trajectory time and throughput rate; however, owing to the nonconvexity of the problem of maximizing the network throughput based on UAV location, we put our effort to achieve these goals by RL approach. Three different approaches have been presented. We first divide the area into a grid and let the UAV explore the environment by flying from point A to point B using both the offline Q-learning and the online SARSA algorithm and the pathloss as the reward. With the intention of maximizing the average payoff, the trajectory in the first scenario is described as a Markov decision process (MDP). According to simulations, MDP produces better results in a smaller area and in less time. In contrast, SARSA performs better in larger environments at the expense of a longer flight duration.