Reinforcement Learning for POMDP: Partitioned Rollout and Policy Iteration with Application to Autonomous Sequential Repair Problems
In this paper we consider infinite horizon discounted dynamic programming\nproblems with finite state and control spaces, and partial state observations.\nWe discuss an algorithm that uses multistep lookahead, truncated rollout with a\nknown base policy, and a terminal cost function approximation. This algorithm\nis also used for policy improvement in an approximate policy iteration scheme,\nwhere successive policies are approximated by using a neural network\nclassifier. A novel feature of our approach is that it is well suited for\ndistributed computation through an extended belief space formulation and the\nuse of a partitioned architecture, which is trained with multiple neural\nnetworks. We apply our methods in simulation to a class of sequential repair\nproblems where a robot inspects and repairs a pipeline with potentially several\nrupture sites under partial information about the state of the pipeline.\n