A Broad-persistent Advising Approach for Deep Interactive Reinforcement Learning in Robotic Environments
Deep Reinforcement Learning (DeepRL) methods have been widely used in\nrobotics to learn about the environment and acquire behaviors autonomously.\nDeep Interactive Reinforcement Learning (DeepIRL) includes interactive feedback\nfrom an external trainer or expert giving advice to help learners choosing\nactions to speed up the learning process. However, current research has been\nlimited to interactions that offer actionable advice to only the current state\nof the agent. Additionally, the information is discarded by the agent after a\nsingle use that causes a duplicate process at the same state for a revisit. In\nthis paper, we present Broad-persistent Advising (BPA), a broad-persistent\nadvising approach that retains and reuses the processed information. It not\nonly helps trainers to give more general advice relevant to similar states\ninstead of only the current state but also allows the agent to speed up the\nlearning process. We test the proposed approach in two continuous robotic\nscenarios, namely, a cart pole balancing task and a simulated robot navigation\ntask. The obtained results show that the performance of the agent using BPA\nimproves while keeping the number of interactions required for the trainer in\ncomparison to the DeepIRL approach.\n