Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images

Understanding the geometric and semantic properties of the scene is crucial in autonomous navigation and particularly challenging in the case of Unmanned Aerial Vehicle (UAV). Such information may be obtained by estimating depth and semantic segmentation maps of the surrounding environment and, for their practical use in autonomous UAV navigation, the procedure must be performed as close to real-time as possible. In this paper, we leverage monocular cameras on aerial robots to predict depth and semantic maps in low-altitude unstructured environments. We propose a joint deep-learning architecture that can perform the two tasks accurately and rapidly, and validate its effectiveness on MidAir and Aeroscapes datasets. Our joint-architecture proved to be competitive or superior to the other single and joint architecture methods while performing its task fast, predicting 20.2 FPS on a single NVIDIA quadro p5000 GPU, with a low memory footprint making it compatible for deployment on hardware. Code is publicly available on this link: https://github.com/Malga-Vision/Co-SemDepth.

Paper

Similar papers

© 2026 NYSGPT2525 LLC