Vision-based perception is of great importance for scene understanding in autonomous systems. RGB-D images are commonly used to capture both semantic and geometric features, but reliable interpretation is challenging due to unavoidable noise in real-world data. In this work, we introduce a diffusion-based framework to address the RGB-D semantic segmentation problem. Additionally, we demonstrate that utilizing a Deformable Attention Transformer as the encoder to extract features from depth images effectively captures the characteristics of invalid regions in depth measurements. Our generative framework shows a greater capacity to model the underlying distribution of RGB-D images, achieving robust performance in challenging scenarios with significantly less training time compared to discriminative methods. Experimental results indicate that our approach achieves State-of-the-Art performance on both the NYUv2 and SUN-RGBD datasets in general and especially in the most challenging of their image data. To demonstrate the practicality of our method, a real-world experiment is conducted to inspect an office and generate its 3D semantic map. Our project page will be available at https://diffusionmms.github.io/
Paper
References (44)
Scroll for more · 32 remaining