Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer

Vision-based perception is of great importance for scene understanding in autonomous systems. RGB-D images are commonly used to capture both semantic and geometric features, but reliable interpretation is challenging due to unavoidable noise in real-world data. In this work, we introduce a diffusion-based framework to address the RGB-D semantic segmentation problem. Additionally, we demonstrate that utilizing a Deformable Attention Transformer as the encoder to extract features from depth images effectively captures the characteristics of invalid regions in depth measurements. Our generative framework shows a greater capacity to model the underlying distribution of RGB-D images, achieving robust performance in challenging scenarios with significantly less training time compared to discriminative methods. Experimental results indicate that our approach achieves State-of-the-Art performance on both the NYUv2 and SUN-RGBD datasets in general and especially in the most challenging of their image data. To demonstrate the practicality of our method, a real-world experiment is conducted to inspect an office and generate its 3D semantic map. Our project page will be available at https://diffusionmms.github.io/

Paper

References (44)

Scroll for more · 32 remaining

Similar papers

© 2026 NYSGPT2525 LLC