Three Ways to Improve Semantic Segmentation with Self-Supervised Depth Estimation

Training deep networks for semantic segmentation requires large amounts of\nlabeled training data, which presents a major challenge in practice, as\nlabeling segmentation masks is a highly labor-intensive process. To address\nthis issue, we present a framework for semi-supervised semantic segmentation,\nwhich is enhanced by self-supervised monocular depth estimation from unlabeled\nimage sequences. In particular, we propose three key contributions: (1) We\ntransfer knowledge from features learned during self-supervised depth\nestimation to semantic segmentation, (2) we implement a strong data\naugmentation by blending images and labels using the geometry of the scene, and\n(3) we utilize the depth feature diversity as well as the level of difficulty\nof learning depth in a student-teacher framework to select the most useful\nsamples to be annotated for semantic segmentation. We validate the proposed\nmodel on the Cityscapes dataset, where all three modules demonstrate\nsignificant performance gains, and we achieve state-of-the-art results for\nsemi-supervised semantic segmentation. The implementation is available at\nhttps://github.com/lhoyer/improving_segmentation_with_selfsupervised_depth.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC