In computational reinforcement learning, a growing body of work seeks to\nexpress an agent's model of the world through predictions about future\nsensations. In this manuscript we focus on predictions expressed as General\nValue Functions: temporally extended estimates of the accumulation of a future\nsignal. One challenge is determining from the infinitely many predictions that\nthe agent could possibly make which might support decision-making. In this\nwork, we contribute a meta-gradient descent method by which an agent can\ndirectly specify what predictions it learns, independent of designer\ninstruction. To that end, we introduce a partially observable domain suited to\nthis investigation. We then demonstrate that through interaction with the\nenvironment an agent can independently select predictions that resolve the\npartial-observability, resulting in performance similar to expertly chosen\nvalue functions. By learning, rather than manually specifying these\npredictions, we enable the agent to identify useful predictions in a\nself-supervised manner, taking a step towards truly autonomous systems.\n
Paper
References (17)
Scroll for more · 5 remaining