We demonstrate how a sampling-based robotic planner can be augmented to learn\nto understand a sequence of natural language commands in a continuous\nconfiguration space to move and manipulate objects. Our approach combines a\ndeep network structured according to the parse of a complex command that\nincludes objects, verbs, spatial relations, and attributes, with a\nsampling-based planner, RRT. A recurrent hierarchical deep network controls how\nthe planner explores the environment, determines when a planned path is likely\nto achieve a goal, and estimates the confidence of each move to trade off\nexploitation and exploration between the network and the planner. Planners are\ndesigned to have near-optimal behavior when information about the task is\nmissing, while networks learn to exploit observations which are available from\nthe environment, making the two naturally complementary. Combining the two\nenables generalization to new maps, new kinds of obstacles, and more complex\nsentences that do not occur in the training set. Little data is required to\ntrain the model despite it jointly acquiring a CNN that extracts features from\nthe environment as it learns the meanings of words. The model provides a level\nof interpretability through the use of attention maps allowing users to see its\nreasoning steps despite being an end-to-end model. This end-to-end model allows\nrobots to learn to follow natural language commands in challenging continuous\nenvironments.\n
Paper
References (29)
Scroll for more · 17 remaining