Receding-Horizon Actor-Critic Design for Learning-Based Control of Nonlinear Continuous-time Systems
Adaptive dynamic programming (ADP) has been recently studied to solve infinite-horizon optimal control problems of nonlinear continuoustime (CT) systems. In this paper, a receding-horizon actor-critic design (RH-ACD) method is proposed to solve the optimal control problem of nonlinear CT systems. In the proposed RH-ACD method, the recedinghorizon control strategy, which is originated from the idea of model predictive control (MPC). The actorcritic structure is designed to approximate the timedependent control policy and value function in each prediction horizon. The network weights of the actor and the critic are updated simultaneously online. The simulation results show that RH-ACD has improved control performance and reduced computational costs when compared with conventional MPC and infinitehorizon ADP.
Paper
Full text
Receding-Horizon Actor-Critic Design for Learning-Based Control of Nonlinear Continuous-time Systems
Semantic Scholar · Engineering · 2020
Abstract
Adaptive dynamic programming (ADP) has been recently studied to solve infinite-horizon optimal control problems of nonlinear continuoustime (CT) systems. In this paper, a receding-horizon actor-critic design (RH-ACD) method is proposed to solve the optimal control problem of nonlinear CT systems. In the proposed RH-ACD method, the recedinghorizon control strategy, which is originated from the idea of model predictive control (MPC). The actorcritic structure is designed to approximate the timedependent control policy and value function in each prediction horizon. The network weights of the actor and the critic are updated simultaneously online. The simulation results show that RH-ACD has improved control performance and reduced computational costs when compared with conventional MPC and infinitehorizon ADP.