We propose a game-theory-based deep-learning tracking control scheme to enable a holonomic flying system to perform an autonomous trajectory tracking task, when considering saturating actuators, adversarial inputs, and nonquadratic cost functionals. The problem is formulated as a two-player zero-sum game, whose online solution is computed by learning the saddle point strategies in real time. Three approximators, namely a critic and two actors, are tuned online using data generated in real time along the system trajectories. The adaptive control character of the algorithm requires a persistence of excitation condition to be a priori validated, which is relaxed by using. a deep-learning approach, based on experience replay with multiple layers. A robustifying control term is added to eliminate the effect of residual errors, leading to asymptotic stability of the equilibrium point of the closed-loop system. A simulation of a target tracking application, where the measurements available to the aerial system are perturbed by persistent adversaries, is performed to validate the effectiveness of the proposed approach.
Paper
Full text
Deep-Learning Tracking for Autonomous Flying Systems Under Adversarial Inputs
Semantic Scholar · Computer Science · 2020
Abstract
We propose a game-theory-based deep-learning tracking control scheme to enable a holonomic flying system to perform an autonomous trajectory tracking task, when considering saturating actuators, adversarial inputs, and nonquadratic cost functionals. The problem is formulated as a two-player zero-sum game, whose online solution is computed by learning the saddle point strategies in real time. Three approximators, namely a critic and two actors, are tuned online using data generated in real time along the system trajectories. The adaptive control character of the algorithm requires a persistence of excitation condition to be a priori validated, which is relaxed by using. a deep-learning approach, based on experience replay with multiple layers. A robustifying control term is added to eliminate the effect of residual errors, leading to asymptotic stability of the equilibrium point of the closed-loop system. A simulation of a target tracking application, where the measurements available to the aerial system are perturbed by persistent adversaries, is performed to validate the effectiveness of the proposed approach.