The majority of learning-based semantic segmentation methods are optimized\nfor daytime scenarios and favorable lighting conditions. Real-world driving\nscenarios, however, entail adverse environmental conditions such as nighttime\nillumination or glare which remain a challenge for existing approaches. In this\nwork, we propose a multimodal semantic segmentation model that can be applied\nduring daytime and nighttime. To this end, besides RGB images, we leverage\nthermal images, making our network significantly more robust. We avoid the\nexpensive annotation of nighttime images by leveraging an existing daytime\nRGB-dataset and propose a teacher-student training approach that transfers the\ndataset's knowledge to the nighttime domain. We further employ a domain\nadaptation method to align the learned feature spaces across the domains and\npropose a novel two-stage training scheme. Furthermore, due to a lack of\nthermal data for autonomous driving, we present a new dataset comprising over\n20,000 time-synchronized and aligned RGB-thermal image pairs. In this context,\nwe also present a novel target-less calibration method that allows for\nautomatic robust extrinsic and intrinsic thermal camera calibration. Among\nothers, we employ our new dataset to show state-of-the-art results for\nnighttime semantic segmentation.\n
Paper
References (36)
Scroll for more · 24 remaining