How to Sense the World: Leveraging Hierarchy in Multimodal Perception for Robust Reinforcement Learning Agents
This work addresses the problem of sensing the world: how to learn a\nmultimodal representation of a reinforcement learning agent's environment that\nallows the execution of tasks under incomplete perceptual conditions. To\naddress such problem, we argue for hierarchy in the design of representation\nmodels and contribute with a novel multimodal representation model, MUSE. The\nproposed model learns hierarchical representations: low-level modality-specific\nrepresentations, encoded from raw observation data, and a high-level multimodal\nrepresentation, encoding joint-modality information to allow robust state\nestimation. We employ MUSE as the sensory representation model of deep\nreinforcement learning agents provided with multimodal observations in Atari\ngames. We perform a comparative study over different designs of reinforcement\nlearning agents, showing that MUSE allows agents to perform tasks under\nincomplete perceptual experience with minimal performance loss. Finally, we\nevaluate the performance of MUSE in literature-standard multimodal scenarios\nwith higher number and more complex modalities, showing that it outperforms\nstate-of-the-art multimodal variational autoencoders in single and\ncross-modality generation.\n
Paper
References (32)
Scroll for more · 20 remaining