Self-Supervised Monocular Depth Estimation: Solving the Dynamic Object Problem by Semantic Guidance
Self-supervised monocular depth estimation presents a powerful method to\nobtain 3D scene information from single camera images, which is trainable on\narbitrary image sequences without requiring depth labels, e.g., from a LiDAR\nsensor. In this work we present a new self-supervised semantically-guided depth\nestimation (SGDepth) method to deal with moving dynamic-class (DC) objects,\nsuch as moving cars and pedestrians, which violate the static-world assumptions\ntypically made during training of such models. Specifically, we propose (i)\nmutually beneficial cross-domain training of (supervised) semantic segmentation\nand self-supervised depth estimation with task-specific network heads, (ii) a\nsemantic masking scheme providing guidance to prevent moving DC objects from\ncontaminating the photometric loss, and (iii) a detection method for frames\nwith non-moving DC objects, from which the depth of DC objects can be learned.\nWe demonstrate the performance of our method on several benchmarks, in\nparticular on the Eigen split, where we exceed all baselines without test-time\nrefinement.\n