Embodied agents operating in human spaces must be able to master how their\nenvironment works: what objects can the agent use, and how can it use them? We\nintroduce a reinforcement learning approach for exploration for interaction,\nwhereby an embodied agent autonomously discovers the affordance landscape of a\nnew unmapped 3D environment (such as an unfamiliar kitchen). Given an\negocentric RGB-D camera and a high-level action space, the agent is rewarded\nfor maximizing successful interactions while simultaneously training an\nimage-based affordance segmentation model. The former yields a policy for\nacting efficiently in new environments to prepare for downstream interaction\ntasks, while the latter yields a convolutional neural network that maps image\nregions to the likelihood they permit each action, densifying the rewards for\nexploration. We demonstrate our idea with AI2-iTHOR. The results show agents\ncan learn how to use new home environments intelligently and that it prepares\nthem to rapidly address various downstream tasks like "find a knife and put it\nin the drawer." Project page:\nhttp://vision.cs.utexas.edu/projects/interaction-exploration/\n
Paper
References (62)
Scroll for more · 38 remaining