Reasoning at multiple levels of temporal abstraction is one of the key\nattributes of intelligence. In reinforcement learning, this is often modeled\nthrough temporally extended courses of actions called options. Options allow\nagents to make predictions and to operate at different levels of abstraction\nwithin an environment. Nevertheless, approaches based on the options framework\noften start with the assumption that a reasonable set of options is known\nbeforehand. When this is not the case, there are no definitive answers for\nwhich options one should consider. In this paper, we argue that the successor\nrepresentation (SR), which encodes states based on the pattern of state\nvisitation that follows them, can be seen as a natural substrate for the\ndiscovery and use of temporal abstractions. To support our claim, we take a big\npicture view of recent results, showing how the SR can be used to discover\noptions that facilitate either temporally-extended exploration or planning. We\ncast these results as instantiations of a general framework for option\ndiscovery in which the agent's representation is used to identify useful\noptions, which are then used to further improve its representation. This\nresults in a virtuous, never-ending, cycle in which both the representation and\nthe options are constantly refined based on each other. Beyond option discovery\nitself, we also discuss how the SR allows us to augment a set of options into a\ncombinatorially large counterpart without additional learning. This is achieved\nthrough the combination of previously learned options. Our empirical evaluation\nfocuses on options discovered for exploration and on the use of the SR to\ncombine them. The results of our experiments shed light on important design\ndecisions involved in the definition of options and demonstrate the synergy of\ndifferent methods based on the SR, such as eigenoptions and the option\nkeyboard.\n
Paper
References (84)
Scroll for more · 38 remaining