Sample-based planning is a powerful family of algorithms for generating\nintelligent behavior from a model of the environment. Generating good candidate\nactions is critical to the success of sample-based planners, particularly in\ncontinuous or large action spaces. Typically, candidate action generation\nexhausts the action space, uses domain knowledge, or more recently, involves\nlearning a stochastic policy to provide such search guidance. In this paper we\nexplore explicitly learning a candidate action generator by optimizing a novel\nobjective, marginal utility. The marginal utility of an action generator\nmeasures the increase in value of an action over previously generated actions.\nWe validate our approach in both curling, a challenging stochastic domain with\ncontinuous state and action spaces, and a location game with a discrete but\nlarge action space. We show that a generator trained with the marginal utility\nobjective outperforms hand-coded schemes built on substantial domain knowledge,\ntrained stochastic policies, and other natural objectives for generating\nactions for sampled-based planners.\n
Paper
References (20)
Scroll for more · 8 remaining