Depth is a vital piece of information for autonomous vehicles to perceive\nobstacles. Due to the relatively low price and small size of monocular cameras,\ndepth estimation from a single RGB image has attracted great interest in the\nresearch community. In recent years, the application of Deep Neural Networks\n(DNNs) has significantly boosted the accuracy of monocular depth estimation\n(MDE). State-of-the-art methods are usually designed on top of complex and\nextremely deep network architectures, which require more computational\nresources and cannot run in real-time without using high-end GPUs. Although\nsome researchers tried to accelerate the running speed, the accuracy of depth\nestimation is degraded because the compressed model does not represent images\nwell. In addition, the inherent characteristic of the feature extractor used by\nthe existing approaches results in severe spatial information loss in the\nproduced feature maps, which also impairs the accuracy of depth estimation on\nsmall sized images. In this study, we are motivated to design a novel and\nefficient Convolutional Neural Network (CNN) that assembles two shallow\nencoder-decoder style subnetworks in succession to address these problems. In\nparticular, we place our emphasis on the trade-off between the accuracy and\nspeed of MDE. Extensive experiments have been conducted on the NYU depth v2,\nKITTI, Make3D and Unreal data sets. Compared with the state-of-the-art\napproaches which have an extremely deep and complex architecture, the proposed\nnetwork not only achieves comparable performance but also runs at a much faster\nspeed on a single, less powerful GPU.\n
Paper
References (66)
Scroll for more · 38 remaining