PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering

We present a new framework for semantic segmentation without annotations via\nclustering. Off-the-shelf clustering methods are limited to curated,\nsingle-label, and object-centric images yet real-world data are dominantly\nuncurated, multi-label, and scene-centric. We extend clustering from images to\npixels and assign separate cluster membership to different instances within\neach image. However, solely relying on pixel-wise feature similarity fails to\nlearn high-level semantic concepts and overfits to low-level visual cues. We\npropose a method to incorporate geometric consistency as an inductive bias to\nlearn invariance and equivariance for photometric and geometric variations.\nWith our novel learning objective, our framework can learn high-level semantic\nconcepts. Our method, PiCIE (Pixel-level feature Clustering using Invariance\nand Equivariance), is the first method capable of segmenting both things and\nstuff categories without any hyperparameter tuning or task-specific\npre-processing. Our method largely outperforms existing baselines on COCO and\nCityscapes with +17.5 Acc. and +4.5 mIoU. We show that PiCIE gives a better\ninitialization for standard supervised training. The code is available at\nhttps://github.com/janghyuncho/PiCIE.\n

Paper

References (67)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC