Environment-agnostic Multitask Learning for Natural Language Grounded Navigation

Recent research efforts enable study for natural language grounded navigation\nin photo-realistic environments, e.g., following natural language instructions\nor dialog. However, existing methods tend to overfit training data in seen\nenvironments and fail to generalize well in previously unseen environments. To\nclose the gap between seen and unseen environments, we aim at learning a\ngeneralized navigation model from two novel perspectives: (1) we introduce a\nmultitask navigation model that can be seamlessly trained on both\nVision-Language Navigation (VLN) and Navigation from Dialog History (NDH)\ntasks, which benefits from richer natural language guidance and effectively\ntransfers knowledge across tasks; (2) we propose to learn environment-agnostic\nrepresentations for the navigation policy that are invariant among the\nenvironments seen during training, thus generalizing better on unseen\nenvironments. Extensive experiments show that environment-agnostic multitask\nlearning significantly reduces the performance gap between seen and unseen\nenvironments, and the navigation agent trained so outperforms baselines on\nunseen environments by 16% (relative measure on success rate) on VLN and 120%\n(goal progress) on NDH. Our submission to the CVDN leaderboard establishes a\nnew state-of-the-art for the NDH task on the holdout test set. Code is\navailable at https://github.com/google-research/valan.\n

Paper

References (53)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC