Traditional transformer-based semantic segmentation re-lies on quantized embeddings. However, our analysis re-veals that autoencoder accuracy on semantic image using quantized embeddings (e.g. VQ- VAE) is 8% lower than continuous-valued embeddings (e.g. KL- VAE). Motivated by this, we propose a continuous-valued embedding framework, CAM-Seg for semantic segmentation. By reformulating semantic image generation as a continuous-valued image-to-embedding diffusion process, our approach eliminates the need for quantized embeddings while preserving fine- grained spatial and semantic details. Our key contribution includes a diffusion-guided autoregressive trans-former that learns a continuous-valued semantic embedding space by modeling long-range dependencies in image features. Our framework contains a unified architecture combining a VAE encoder for continuous-valued feature extraction, a diffusion-guided transformer for conditioned embedding generation, and a VAE decoder for semantic image reconstruction. Our setting facilitates zero-shot domain adaptation capabilities enabled by the continuity of the embedding space. Experiments across diverse datasets (e.g., Cityscapes and domain-shifted variants) demonstrate state-of-the-art robustness to distribution shifts, including adverse weather (e.g., fog, snow) and viewpoint variations. Our model also exhibits strong noise resilience, achieving robust performance $(\approx 95\% AP$ compared to baseline) under gaussian noise, moderate motion blur, and moderate brightness/contrast variations, while experiencing only a moderate impact $(\approx 90\%$ AP compared to baseline) from 50% salt and pepper noise, saturation and hue shifts. [We will release the code.]
Paper
References (50)
Scroll for more · 38 remaining