Many reinforcement learning (RL) environments consist of independent entities\nthat interact sparsely. In such environments, RL agents have only limited\ninfluence over other entities in any particular situation. Our idea in this\nwork is that learning can be efficiently guided by knowing when and what the\nagent can influence with its actions. To achieve this, we introduce a measure\nof \\emph{situation-dependent causal influence} based on conditional mutual\ninformation and show that it can reliably detect states of influence. We then\npropose several ways to integrate this measure into RL algorithms to improve\nexploration and off-policy learning. All modified algorithms show strong\nincreases in data efficiency on robotic manipulation tasks.\n
Paper
References (70)
Scroll for more · 38 remaining