Railroad is not a Train: Saliency as Pseudo-pixel Supervision for Weakly Supervised Semantic Segmentation
Existing studies in weakly-supervised semantic segmentation (WSSS) using\nimage-level weak supervision have several limitations: sparse object coverage,\ninaccurate object boundaries, and co-occurring pixels from non-target objects.\nTo overcome these challenges, we propose a novel framework, namely Explicit\nPseudo-pixel Supervision (EPS), which learns from pixel-level feedback by\ncombining two weak supervisions; the image-level label provides the object\nidentity via the localization map and the saliency map from the off-the-shelf\nsaliency detection model offers rich boundaries. We devise a joint training\nstrategy to fully utilize the complementary relationship between both\ninformation. Our method can obtain accurate object boundaries and discard\nco-occurring pixels, thereby significantly improving the quality of\npseudo-masks. Experimental results show that the proposed method remarkably\noutperforms existing methods by resolving key challenges of WSSS and achieves\nthe new state-of-the-art performance on both PASCAL VOC 2012 and MS COCO 2014\ndatasets.\n
Paper
References (52)
Scroll for more · 38 remaining