Active Trajectory Estimation for Partially Observed Markov Decision\n Processes via Conditional Entropy

In this paper, we consider the problem of controlling a partially observed\nMarkov decision process (POMDP) in order to actively estimate its state\ntrajectory over a fixed horizon with minimal uncertainty. We pose a novel\nactive smoothing problem in which the objective is to directly minimise the\nsmoother entropy, that is, the conditional entropy of the (joint) state\ntrajectory distribution of concern in fixed-interval Bayesian smoothing. Our\nformulation contrasts with prior active approaches that minimise the sum of\nconditional entropies of the (marginal) state estimates provided by Bayesian\nfilters. By establishing a novel form of the smoother entropy in terms of the\nPOMDP belief (or information) state, we show that our active smoothing problem\ncan be reformulated as a (fully observed) Markov decision process with a value\nfunction that is concave in the belief state. The concavity of the value\nfunction is of particular importance since it enables the approximate solution\nof our active smoothing problem using piecewise-linear function approximations\nin conjunction with standard POMDP solvers. We illustrate the approximate\nsolution of our active smoothing problem in simulation and compare its\nperformance to alternative approaches based on minimising marginal state\nestimate uncertainties.\n

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

© 2026 NYSGPT2525 LLC