Multi-view stereo (MVS) is the golden mean between the accuracy of active\ndepth sensing and the practicality of monocular depth estimation. Cost volume\nbased approaches employing 3D convolutional neural networks (CNNs) have\nconsiderably improved the accuracy of MVS systems. However, this accuracy comes\nat a high computational cost which impedes practical adoption. Distinct from\ncost volume approaches, we propose an efficient depth estimation approach by\nfirst (a) detecting and evaluating descriptors for interest points, then (b)\nlearning to match and triangulate a small set of interest points, and finally\n(c) densifying this sparse set of 3D points using CNNs. An end-to-end network\nefficiently performs all three steps within a deep learning framework and\ntrained with intermediate 2D image and 3D geometric supervision, along with\ndepth supervision. Crucially, our first step complements pose estimation using\ninterest point detection and descriptor learning. We demonstrate\nstate-of-the-art results on depth estimation with lower compute for different\nscene lengths. Furthermore, our method generalizes to newer environments and\nthe descriptors output by our network compare favorably to strong baselines.\nCode is available at https://github.com/magicleap/DELTAS\n
Paper
References (49)
Scroll for more · 37 remaining