SENTINEL: Taming Uncertainty with Ensemble-based Distributional Reinforcement Learning

In this paper, we consider risk-sensitive sequential decision-making in\nReinforcement Learning (RL). Our contributions are two-fold. First, we\nintroduce a novel and coherent quantification of risk, namely composite risk,\nwhich quantifies the joint effect of aleatory and epistemic risk during the\nlearning process. Existing works considered either aleatory or epistemic risk\nindividually, or as an additive combination. We prove that the additive\nformulation is a particular case of the composite risk when the epistemic risk\nmeasure is replaced with expectation. Thus, the composite risk is more\nsensitive to both aleatory and epistemic uncertainty than the individual and\nadditive formulations. We also propose an algorithm, SENTINEL-K, based on\nensemble bootstrapping and distributional RL for representing epistemic and\naleatory uncertainty respectively. The ensemble of K learners uses Follow The\nRegularised Leader (FTRL) to aggregate the return distributions and obtain the\ncomposite risk. We experimentally verify that SENTINEL-K estimates the return\ndistribution better, and while used with composite risk estimates, demonstrates\nhigher risk-sensitive performance than state-of-the-art risk-sensitive and\ndistributional RL algorithms.\n

Paper

References (52)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC