Reinforcement learning has shown strong potential in learning optimal control strategy by modelling policy and/or value function. Even though policy and value function forms duality regarding the Bellman equation, there is no structure unifies this two branches. In this paper, we propose to use an convex optimization layer to combine these two branches which enables universal compatibility with all reinforcement learning algorithm without modification of the model structure. Design and training issues will be explained and validated by both linear and nonlinear control.
Paper
References (21)
Scroll for more · 9 remaining