Monocular Depth Estimation through Virtual-world Supervision and Real-world SfM Self-Supervision

Depth information is essential for on-board perception in autonomous driving\nand driver assistance. Monocular depth estimation (MDE) is very appealing since\nit allows for appearance and depth being on direct pixelwise correspondence\nwithout further calibration. Best MDE models are based on Convolutional Neural\nNetworks (CNNs) trained in a supervised manner, i.e., assuming pixelwise ground\ntruth (GT). Usually, this GT is acquired at training time through a calibrated\nmulti-modal suite of sensors. However, also using only a monocular system at\ntraining time is cheaper and more scalable. This is possible by relying on\nstructure-from-motion (SfM) principles to generate self-supervision.\nNevertheless, problems of camouflaged objects, visibility changes,\nstatic-camera intervals, textureless areas, and scale ambiguity, diminish the\nusefulness of such self-supervision. In this paper, we perform monocular depth\nestimation by virtual-world supervision (MonoDEVS) and real-world SfM\nself-supervision. We compensate the SfM self-supervision limitations by\nleveraging virtual-world images with accurate semantic and depth supervision\nand addressing the virtual-to-real domain gap. Our MonoDEVSNet outperforms\nprevious MDE CNNs trained on monocular and even stereo sequences.\n

Paper

References (70)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC