Traditional policy gradient methods are fundamentally flawed. Natural gradients converge quicker and better, forming the foundation of contemporary Reinforcement Learning such as Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO). This lecture note aims to clarify the intuition behind natural policy gradients, focusing on the thought process and the key mathematical constructs.
Paper
References (8)
05In terms of lecture slides, I found the following ones particularly helpful. • Levine, S. Advanced Policy Gradients (CS 285)
06Natural Gradient Descent2018
07For the origins of Natural Policy gradients, I would suggest reading the foundational papers by Amari (1998) and Kokade (2001), as well as the more recent reflection by Martens2020
08Finally, the following posts provide great explanations from different angles