Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes
In the theory of Partially Observed Markov Decision Processes (POMDPs),\nexistence of optimal policies have in general been established via converting\nthe original partially observed stochastic control problem to a fully observed\none on the belief space, leading to a belief-MDP. However, computing an optimal\npolicy for this fully observed model, and so for the original POMDP, using\nclassical dynamic or linear programming methods is challenging even if the\noriginal system has finite state and action spaces, since the state space of\nthe fully observed belief-MDP model is always uncountable. Furthermore, there\nexist very few rigorous value function approximation and optimal policy\napproximation results, as regularity conditions needed often require a tedious\nstudy involving the spaces of probability measures leading to properties such\nas Feller continuity. In this paper, we study a planning problem for POMDPs\nwhere the system dynamics and measurement channel model are assumed to be\nknown. We construct an approximate belief model by discretizing the belief\nspace using only finite window information variables. We then find optimal\npolicies for the approximate model and we rigorously establish near optimality\nof the constructed finite window control policies in POMDPs under mild\nnon-linear filter stability conditions and the assumption that the measurement\nand action sets are finite (and the state space is real vector valued). We also\nestablish a rate of convergence result which relates the finite window memory\nsize and the approximation error bound, where the rate of convergence is\nexponential under explicit and testable exponential filter stability\nconditions. While there exist many experimental results and few rigorous\nasymptotic convergence results, an explicit rate of convergence result is new\nin the literature, to our knowledge.\n
Paper
References (44)
Scroll for more · 32 remaining