Reinforcement Learning Using Expectation Maximization Based Guided\n Policy Search for Stochastic Dynamics
Guided policy search algorithms have been proven to work with incredible\naccuracy for not only controlling a complicated dynamical system, but also\nlearning optimal policies from various unseen instances. One assumes true\nnature of the states in almost all of the well known policy search and learning\nalgorithms. This paper deals with a trajectory optimization procedure for an\nunknown dynamical system subject to measurement noise using expectation\nmaximization and extends it to learning (optimal) policies which have less\nnoise because of lower variance in the optimal trajectories. Theoretical and\nempirical evidence of learnt optimal policies of the new approach is depicted\nin comparison to some well known baselines which are evaluated on an autonomous\nsystem with widely used performance metrics.\n