The Reinforce Policy Gradient Algorithm Revisited

We revisit the Reinforce policy gradient algorithm that works with full cost returns obtained over random length episodes. We propose a new Reinforce type algorithm that estimates the policy gradient using a function measurement over a perturbed parameter using a smoothed functional based gradient estimator. We observe that even though we estimate the gradient of the performance objective using sample performance (and not the sample gradient), the algorithm converges to a neighborhood of a local minimum. We further describe the main convergence result.

Paper

Similar papers

© 2026 NYSGPT2525 LLC