RidgeSfM: Structure from Motion via Robust Pairwise Matching Under Depth Uncertainty

We consider the problem of simultaneously estimating a dense depth map and\ncamera pose for a large set of images of an indoor scene. While classical SfM\npipelines rely on a two-step approach where cameras are first estimated using a\nbundle adjustment in order to ground the ensuing multi-view stereo stage, both\nour poses and dense reconstructions are a direct output of an altered bundle\nadjuster. To this end, we parametrize each depth map with a linear combination\nof a limited number of basis "depth-planes" predicted in a monocular fashion by\na deep net. Using a set of high-quality sparse keypoint matches, we optimize\nover the per-frame linear combinations of depth planes and camera poses to form\na geometrically consistent cloud of keypoints. Although our bundle adjustment\nonly considers sparse keypoints, the inferred linear coefficients of the basis\nplanes immediately give us dense depth maps. RidgeSfM is able to collectively\nalign hundreds of frames, which is its main advantage over recent memory-heavy\ndeep alternatives that can align at most 10 frames. Quantitative comparisons\nreveal performance superior to a state-of-the-art large-scale SfM pipeline.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC