Active hypothesis testing in unknown environments using recurrent neural networks and model free reinforcement learning
A combination of deep reinforcement learning and supervised learning is proposed for active sequential hypothesis testing in completely unknown environments. We make no assumptions about the prior probability, the action and observation sets, as well as the observation generating process. Experiments with synthetic Bernoulli and real cybersecurity data demonstrate that our method performs competitively and sometimes better than the Chernoff test, in both finite and infinite horizon problems.