Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach

We propose a unified framework to study policy evaluation (PE) and the\nassociated temporal difference (TD) methods for reinforcement learning in\ncontinuous time and space. We show that PE is equivalent to maintaining the\nmartingale condition of a process. From this perspective, we find that the\nmean--square TD error approximates the quadratic variation of the martingale\nand thus is not a suitable objective for PE. We present two methods to use the\nmartingale characterization for designing PE algorithms. The first one\nminimizes a "martingale loss function", whose solution is proved to be the best\napproximation of the true value function in the mean--square sense. This method\ninterprets the classical gradient Monte-Carlo algorithm. The second method is\nbased on a system of equations called the "martingale orthogonality conditions"\nwith test functions. Solving these equations in different ways recovers various\nclassical TD algorithms, such as TD($\\lambda$), LSTD, and GTD. Different\nchoices of test functions determine in what sense the resulting solutions\napproximate the true value function. Moreover, we prove that any convergent\ntime-discretized algorithm converges to its continuous-time counterpart as the\nmesh size goes to zero, and we provide the convergence rate. We demonstrate the\ntheoretical results and corresponding algorithms with numerical experiments and\napplications.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC