Monocular visual odometry (VO) has attracted extensive research attention by\nproviding real-time vehicle motion from cost-effective camera images. However,\nstate-of-the-art optimization-based monocular VO methods suffer from the scale\ninconsistency problem for long-term predictions. Deep learning has recently\nbeen introduced to address this issue by leveraging stereo sequences or\nground-truth motions in the training dataset. However, it comes at an\nadditional cost for data collection, and such training data may not be\navailable in all datasets. In this work, we propose VRVO, a novel framework for\nretrieving the absolute scale from virtual data that can be easily obtained\nfrom modern simulation environments, whereas in the real domain no stereo or\nground-truth data are required in either the training or inference phases.\nSpecifically, we first train a scale-aware disparity network using both\nmonocular real images and stereo virtual data. The virtual-to-real domain gap\nis bridged by using an adversarial training strategy to map images from both\ndomains into a shared feature space. The resulting scale-consistent disparities\nare then integrated with a direct VO system by constructing a virtual stereo\nobjective that ensures the scale consistency over long trajectories.\nAdditionally, to address the suboptimality issue caused by the separate\noptimization backend and the learning process, we further propose a mutual\nreinforcement pipeline that allows bidirectional information flow between\nlearning and optimization, which boosts the robustness and accuracy of each\nother. We demonstrate the effectiveness of our framework on the KITTI and\nvKITTI2 datasets.\n
Paper
References (40)
Scroll for more · 28 remaining