DeFeat-Net: General Monocular Depth via Simultaneous Unsupervised Representation Learning

In the current monocular depth research, the dominant approach is to employ\nunsupervised training on large datasets, driven by warped photometric\nconsistency. Such approaches lack robustness and are unable to generalize to\nchallenging domains such as nighttime scenes or adverse weather conditions\nwhere assumptions about photometric consistency break down.\n We propose DeFeat-Net (Depth & Feature network), an approach to\nsimultaneously learn a cross-domain dense feature representation, alongside a\nrobust depth-estimation framework based on warped feature consistency. The\nresulting feature representation is learned in an unsupervised manner with no\nexplicit ground-truth correspondences required.\n We show that within a single domain, our technique is comparable to both the\ncurrent state of the art in monocular depth estimation and supervised feature\nrepresentation learning. However, by simultaneously learning features, depth\nand motion, our technique is able to generalize to challenging domains,\nallowing DeFeat-Net to outperform the current state-of-the-art with around 10%\nreduction in all error measures on more challenging sequences such as nighttime\ndriving.\n

Paper

References (84)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC