Self-supervised Monocular Trained Depth Estimation using Self-attention and Discrete Disparity Volume
Monocular depth estimation has become one of the most studied applications in\ncomputer vision, where the most accurate approaches are based on fully\nsupervised learning models. However, the acquisition of accurate and large\nground truth data sets to model these fully supervised methods is a major\nchallenge for the further development of the area. Self-supervised methods\ntrained with monocular videos constitute one the most promising approaches to\nmitigate the challenge mentioned above due to the wide-spread availability of\ntraining data. Consequently, they have been intensively studied, where the main\nideas explored consist of different types of model architectures, loss\nfunctions, and occlusion masks to address non-rigid motion. In this paper, we\npropose two new ideas to improve self-supervised monocular trained depth\nestimation: 1) self-attention, and 2) discrete disparity prediction. Compared\nwith the usual localised convolution operation, self-attention can explore a\nmore general contextual information that allows the inference of similar\ndisparity values at non-contiguous regions of the image. Discrete disparity\nprediction has been shown by fully supervised methods to provide a more robust\nand sharper depth estimation than the more common continuous disparity\nprediction, besides enabling the estimation of depth uncertainty. We show that\nthe extension of the state-of-the-art self-supervised monocular trained depth\nestimator Monodepth2 with these two ideas allows us to design a model that\nproduces the best results in the field in KITTI 2015 and Make3D, closing the\ngap with respect self-supervised stereo training and fully supervised\napproaches.\n