JueWu-MC: Playing Minecraft with Sample-efficient Hierarchical Reinforcement Learning

Learning rational behaviors in open-world games like Minecraft remains to be\nchallenging for Reinforcement Learning (RL) research due to the compound\nchallenge of partial observability, high-dimensional visual perception and\ndelayed reward. To address this, we propose JueWu-MC, a sample-efficient\nhierarchical RL approach equipped with representation learning and imitation\nlearning to deal with perception and exploration. Specifically, our approach\nincludes two levels of hierarchy, where the high-level controller learns a\npolicy to control over options and the low-level workers learn to solve each\nsub-task. To boost the learning of sub-tasks, we propose a combination of\ntechniques including 1) action-aware representation learning which captures\nunderlying relations between action and representation, 2) discriminator-based\nself-imitation learning for efficient exploration, and 3) ensemble behavior\ncloning with consistency filtering for policy robustness. Extensive experiments\nshow that JueWu-MC significantly improves sample efficiency and outperforms a\nset of baselines by a large margin. Notably, we won the championship of the\nNeurIPS MineRL 2021 research competition and achieved the highest performance\nscore ever.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC