Distributional Multivariate Policy Evaluation and Exploration with the Bellman GAN

The recently proposed distributional approach to reinforcement learning\n(DiRL) is centered on learning the distribution of the reward-to-go, often\nreferred to as the value distribution. In this work, we show that the\ndistributional Bellman equation, which drives DiRL methods, is equivalent to a\ngenerative adversarial network (GAN) model. In this formulation, DiRL can be\nseen as learning a deep generative model of the value distribution, driven by\nthe discrepancy between the distribution of the current value, and the\ndistribution of the sum of current reward and next value. We use this insight\nto propose a GAN-based approach to DiRL, which leverages the strengths of GANs\nin learning distributions of high-dimensional data. In particular, we show that\nour GAN approach can be used for DiRL with multivariate rewards, an important\nsetting which cannot be tackled with prior methods. The multivariate setting\nalso allows us to unify learning the distribution of values and state\ntransitions, and we exploit this idea to devise a novel exploration method that\nis driven by the discrepancy in estimating both values and states.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC