Learning a Geometric Representation for Data-Efficient Depth Estimation via Gradient Field and Contrastive Loss

Estimating a depth map from a single RGB image has been investigated widely\nfor localization, mapping, and 3-dimensional object detection. Recent studies\non a single-view depth estimation are mostly based on deep Convolutional neural\nNetworks (ConvNets) which require a large amount of training data paired with\ndensely annotated labels. Depth annotation tasks are both expensive and\ninefficient, so it is inevitable to leverage RGB images which can be collected\nvery easily to boost the performance of ConvNets without depth labels. However,\nmost self-supervised learning algorithms are focused on capturing the semantic\ninformation of images to improve the performance in classification or object\ndetection, not in depth estimation. In this paper, we show that existing\nself-supervised methods do not perform well on depth estimation and propose a\ngradient-based self-supervised learning algorithm with momentum contrastive\nloss to help ConvNets extract the geometric information with unlabeled images.\nAs a result, the network can estimate the depth map accurately with a\nrelatively small amount of annotated data. To show that our method is\nindependent of the model structure, we evaluate our method with two different\nmonocular depth estimation algorithms. Our method outperforms the previous\nstate-of-the-art self-supervised learning algorithms and shows the efficiency\nof labeled data in triple compared to random initialization on the NYU Depth v2\ndataset.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC