How to Train PointGoal Navigation Agents on a (Sample and Compute) Budget

PointGoal navigation has seen significant recent interest and progress,\nspurred on by the Habitat platform and associated challenge. In this paper, we\nstudy PointGoal navigation under both a sample budget (75 million frames) and a\ncompute budget (1 GPU for 1 day). We conduct an extensive set of experiments,\ncumulatively totaling over 50,000 GPU-hours, that let us identify and discuss a\nnumber of ostensibly minor but significant design choices -- the advantage\nestimation procedure (a key component in training), visual encoder\narchitecture, and a seemingly minor hyper-parameter change. Overall, these\ndesign choices to lead considerable and consistent improvements over the\nbaselines present in Savva et al. Under a sample budget, performance for RGB-D\nagents improves 8 SPL on Gibson (14% relative improvement) and 20 SPL on\nMatterport3D (38% relative improvement). Under a compute budget, performance\nfor RGB-D agents improves by 19 SPL on Gibson (32% relative improvement) and 35\nSPL on Matterport3D (220% relative improvement). We hope our findings and\nrecommendations will make serve to make the community's experiments more\nefficient.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC