UnRectDepthNet: Self-Supervised Monocular Depth Estimation using a Generic Framework for Handling Common Camera Distortion Models

In classical computer vision, rectification is an integral part of multi-view\ndepth estimation. It typically includes epipolar rectification and lens\ndistortion correction. This process simplifies the depth estimation\nsignificantly, and thus it has been adopted in CNN approaches. However,\nrectification has several side effects, including a reduced field of view\n(FOV), resampling distortion, and sensitivity to calibration errors. The\neffects are particularly pronounced in case of significant distortion (e.g.,\nwide-angle fisheye cameras). In this paper, we propose a generic scale-aware\nself-supervised pipeline for estimating depth, euclidean distance, and visual\nodometry from unrectified monocular videos. We demonstrate a similar level of\nprecision on the unrectified KITTI dataset with barrel distortion comparable to\nthe rectified KITTI dataset. The intuition being that the rectification step\ncan be implicitly absorbed within the CNN model, which learns the distortion\nmodel without increasing complexity. Our approach does not suffer from a\nreduced field of view and avoids computational costs for rectification at\ninference time. To further illustrate the general applicability of the proposed\nframework, we apply it to wide-angle fisheye cameras with 190$^\\circ$\nhorizontal field of view. The training framework UnRectDepthNet takes in the\ncamera distortion model as an argument and adapts projection and unprojection\nfunctions accordingly. The proposed algorithm is evaluated further on the KITTI\nrectified dataset, and we achieve state-of-the-art results that improve upon\nour previous work FisheyeDistanceNet. Qualitative results on a distorted test\nscene video sequence indicate excellent performance\nhttps://youtu.be/K6pbx3bU4Ss.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC