Hierarchical Reinforcement Learning (HRL) algorithms such as Option-Critic (OC) and Multi-updates Option Critic (MOC) advance the learning of reusable options but struggle in multi-goal environments with sparse rewards. We propose MOC-HER, which integrates the Hindsight Experience Replay (HER) mechanism into MOC. By relabeling goals from achieved outcomes, MOC-HER addresses sparse reward environments that are intractable for the original MOC. However, for object manipulation tasks, where rewards are determined by object placement rather than agent-centric states, standard relabeling is often insufficient. We introduce Dual Objectives Hindsight Experience Replay (2HER), which augments goal relabeling with virtual goals from the agent's effector positions. This encourages both effective object interaction and task completion. In robotic manipulation tasks, MOC-2HER achieves success rates up to 90%, compared to under 11% for MOC and MOC-HER.
Paper
References (29)
Scroll for more · 17 remaining