Depth estimation and semantic segmentation play essential roles in scene\nunderstanding. The state-of-the-art methods employ multi-task learning to\nsimultaneously learn models for these two tasks at the pixel-wise level. They\nusually focus on sharing the common features or stitching feature maps from the\ncorresponding branches. However, these methods lack in-depth consideration on\nthe correlation of the geometric cues and the scene parsing. In this paper, we\nfirst introduce the concept of semantic objectness to exploit the geometric\nrelationship of these two tasks through an analysis of the imaging process,\nthen propose a Semantic Object Segmentation and Depth Estimation Network\n(SOSD-Net) based on the objectness assumption. To the best of our knowledge,\nSOSD-Net is the first network that exploits the geometry constraint for\nsimultaneous monocular depth estimation and semantic segmentation. In addition,\nconsidering the mutual implicit relationship between these two tasks, we\nexploit the iterative idea from the expectation-maximization algorithm to train\nthe proposed network more effectively. Extensive experimental results on the\nCityscapes and NYU v2 dataset are presented to demonstrate the superior\nperformance of the proposed approach.\n
Paper
References (67)
Scroll for more · 38 remaining