SynDistNet: Self-Supervised Monocular Fisheye Camera Distance Estimation Synergized with Semantic Segmentation for Autonomous Driving

State-of-the-art self-supervised learning approaches for monocular depth\nestimation usually suffer from scale ambiguity. They do not generalize well\nwhen applied on distance estimation for complex projection models such as in\nfisheye and omnidirectional cameras. This paper introduces a novel multi-task\nlearning strategy to improve self-supervised monocular distance estimation on\nfisheye and pinhole camera images. Our contribution to this work is threefold:\nFirstly, we introduce a novel distance estimation network architecture using a\nself-attention based encoder coupled with robust semantic feature guidance to\nthe decoder that can be trained in a one-stage fashion. Secondly, we integrate\na generalized robust loss function, which improves performance significantly\nwhile removing the need for hyperparameter tuning with the reprojection loss.\nFinally, we reduce the artifacts caused by dynamic objects violating static\nworld assumptions using a semantic masking strategy. We significantly improve\nupon the RMSE of previous work on fisheye by 25% reduction in RMSE. As there is\nlittle work on fisheye cameras, we evaluated the proposed method on KITTI using\na pinhole model. We achieved state-of-the-art performance among self-supervised\nmethods without requiring an external scale estimation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC