UAV-based Mobile Acoustic Sink for Dynamic Communication Coverage via Multi-Agent Reinforcement Learning Approach

Unmanned aerial vehicles (UAVs) are considered as promising devices for intelligent mission execution in the smart ocean due to their flexibility, mobility, and ability to deploy as base stations. In this paper, we aim to design a multi-agent deep reinforcement learning-based control solution that uses UAVs as mobile acoustic sinks to provide optimal communication coverage for autonomous underwater vehicles (AUVs) to monitor complex conditions in the ocean. In order to solve the problem of cross-boundary communication between UAVs and AUVs, we establish real-time communication with the AUVs by equipping the UAVs with hydrophones that transmit the data collected by the AUVs. The goal is to address the coverage imbalance issue caused by the movement of the swarm of AUVs. We model the problem as a partially observable markov decision process (POMDP), and then design a reasonable reward function based on local information. Based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm, we propose a dynamic deployment algorithm to optimize the dynamic deployment process of the UAVs in order to achieve the maximum communication coverage.

Paper

Full text

PDF

UAV-based Mobile Acoustic Sink for Dynamic Communication Coverage via Multi-Agent Reinforcement Learning Approach

Semantic Scholar · Engineering · 2023

Abstract

Unmanned aerial vehicles (UAVs) are considered as promising devices for intelligent mission execution in the smart ocean due to their flexibility, mobility, and ability to deploy as base stations. In this paper, we aim to design a multi-agent deep reinforcement learning-based control solution that uses UAVs as mobile acoustic sinks to provide optimal communication coverage for autonomous underwater vehicles (AUVs) to monitor complex conditions in the ocean. In order to solve the problem of cross-boundary communication between UAVs and AUVs, we establish real-time communication with the AUVs by equipping the UAVs with hydrophones that transmit the data collected by the AUVs. The goal is to address the coverage imbalance issue caused by the movement of the swarm of AUVs. We model the problem as a partially observable markov decision process (POMDP), and then design a reasonable reward function based on local information. Based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm, we propose a dynamic deployment algorithm to optimize the dynamic deployment process of the UAVs in order to achieve the maximum communication coverage.

References (7)

Similar papers

© 2026 NYSGPT2525 LLC