Pix2Point: Learning Outdoor 3D Using Sparse Point Clouds and Optimal Transport

Good quality reconstruction and comprehension of a scene rely on 3D\nestimation methods. The 3D information was usually obtained from images by\nstereo-photogrammetry, but deep learning has recently provided us with\nexcellent results for monocular depth estimation. Building up a sufficiently\nlarge and rich training dataset to achieve these results requires onerous\nprocessing. In this paper, we address the problem of learning outdoor 3D point\ncloud from monocular data using a sparse ground-truth dataset. We propose\nPix2Point, a deep learning-based approach for monocular 3D point cloud\nprediction, able to deal with complete and challenging outdoor scenes. Our\nmethod relies on a 2D-3D hybrid neural network architecture, and a supervised\nend-to-end minimisation of an optimal transport divergence between point\nclouds. We show that, when trained on sparse point clouds, our simple promising\napproach achieves a better coverage of 3D outdoor scenes than efficient\nmonocular depth methods.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC