Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

Vision-and-language navigation (VLN) is a task in which an agent is embodied\nin a realistic 3D environment and follows an instruction to reach the goal\nnode. While most of the previous studies have built and investigated a\ndiscriminative approach, we notice that there are in fact two possible\napproaches to building such a VLN agent: discriminative \\textit{and}\ngenerative. In this paper, we design and investigate a generative\nlanguage-grounded policy which uses a language model to compute the\ndistribution over all possible instructions i.e. all possible sequences of\nvocabulary tokens given action and the transition history. In experiments, we\nshow that the proposed generative approach outperforms the discriminative\napproach in the Room-2-Room (R2R) and Room-4-Room (R4R) datasets, especially in\nthe unseen environments. We further show that the combination of the generative\nand discriminative policies achieves close to the state-of-the art results in\nthe R2R dataset, demonstrating that the generative and discriminative policies\ncapture the different aspects of VLN.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC