The vision-language navigation (VLN) task requires an agent to reach a target\nwith the guidance of natural language instruction. Previous works learn to\nnavigate step-by-step following an instruction. However, these works may fail\nto discriminate the similarities and discrepancies across\ninstruction-trajectory pairs and ignore the temporal continuity of\nsub-instructions. These problems hinder agents from learning distinctive\nvision-and-language representations, harming the robustness and\ngeneralizability of the navigation policy. In this paper, we propose a\nContrastive Instruction-Trajectory Learning (CITL) framework that explores\ninvariance across similar data samples and variance across different ones to\nlearn distinctive representations for robust navigation. Specifically, we\npropose: (1) a coarse-grained contrastive learning objective to enhance\nvision-and-language representations by contrasting semantics of full trajectory\nobservations and instructions, respectively; (2) a fine-grained contrastive\nlearning objective to perceive instructions by leveraging the temporal\ninformation of the sub-instructions; (3) a pairwise sample-reweighting\nmechanism for contrastive learning to mine hard samples and hence mitigate the\ninfluence of data sampling bias in contrastive learning. Our CITL can be easily\nintegrated with VLN backbones to form a new learning paradigm and achieve\nbetter generalizability in unseen environments. Extensive experiments show that\nthe model with CITL surpasses the previous state-of-the-art methods on R2R,\nR4R, and RxR.\n
Paper
References (65)
Scroll for more · 38 remaining