Learning good feature representations is important for deep reinforcement\nlearning (RL). However, with limited experience, RL often suffers from data\ninefficiency for training. For un-experienced or less-experienced trajectories\n(i.e., state-action sequences), the lack of data limits the use of them for\nbetter feature learning. In this work, we propose a novel method, dubbed\nPlayVirtual, which augments cycle-consistent virtual trajectories to enhance\nthe data efficiency for RL feature representation learning. Specifically,\nPlayVirtual predicts future states in the latent space based on the current\nstate and action by a dynamics model and then predicts the previous states by a\nbackward dynamics model, which forms a trajectory cycle. Based on this, we\naugment the actions to generate a large amount of virtual state-action\ntrajectories. Being free of groudtruth state supervision, we enforce a\ntrajectory to meet the cycle consistency constraint, which can significantly\nenhance the data efficiency. We validate the effectiveness of our designs on\nthe Atari and DeepMind Control Suite benchmarks. Our method achieves the\nstate-of-the-art performance on both benchmarks.\n
Paper
References (70)
Scroll for more · 38 remaining