In this paper, we consider risk-sensitive sequential decision-making in\nReinforcement Learning (RL). Our contributions are two-fold. First, we\nintroduce a novel and coherent quantification of risk, namely composite risk,\nwhich quantifies the joint effect of aleatory and epistemic risk during the\nlearning process. Existing works considered either aleatory or epistemic risk\nindividually, or as an additive combination. We prove that the additive\nformulation is a particular case of the composite risk when the epistemic risk\nmeasure is replaced with expectation. Thus, the composite risk is more\nsensitive to both aleatory and epistemic uncertainty than the individual and\nadditive formulations. We also propose an algorithm, SENTINEL-K, based on\nensemble bootstrapping and distributional RL for representing epistemic and\naleatory uncertainty respectively. The ensemble of K learners uses Follow The\nRegularised Leader (FTRL) to aggregate the return distributions and obtain the\ncomposite risk. We experimentally verify that SENTINEL-K estimates the return\ndistribution better, and while used with composite risk estimates, demonstrates\nhigher risk-sensitive performance than state-of-the-art risk-sensitive and\ndistributional RL algorithms.\n
Paper
References (52)
Scroll for more · 38 remaining