Methods for learning and planning in sequential decision problems often assume the learner is fully aware of all possible states and actions in advance. This assumption is sometimes untenable: evidence gathered via domain exploration or external advice may reveal not just information about which of the currently known states are probable, but that entirely new states or actions are possible. This paper provides a model-based method for learning factored markov decision problems from both domain exploration and contextually relevant expert corrections in a way which guarantees convergence to near-optimal behaviour, even when the agent is initially unaware of actions or belief variables that are critical to achieving success. Our experiments show that our agent converges quickly on the optimal policy for both large and small decision problems. We also explore how an expert's tolerance towards the agent's mistakes affects the agent's ability to achieve optimal behaviour.
Paper
References (36)
Scroll for more · 24 remaining