Find it if You Can: End-to-End Adversarial Erasing for Weakly-Supervised Semantic Segmentation
Semantic segmentation is a task that traditionally requires a large dataset\nof pixel-level ground truth labels, which is time-consuming and expensive to\nobtain. Recent advancements in the weakly-supervised setting show that\nreasonable performance can be obtained by using only image-level labels.\nClassification is often used as a proxy task to train a deep neural network\nfrom which attention maps are extracted. However, the classification task needs\nonly the minimum evidence to make predictions, hence it focuses on the most\ndiscriminative object regions. To overcome this problem, we propose a novel\nformulation of adversarial erasing of the attention maps. In contrast to\nprevious adversarial erasing methods, we optimize two networks with opposing\nloss functions, which eliminates the requirement of certain suboptimal\nstrategies; for instance, having multiple training steps that complicate the\ntraining process or a weight sharing policy between networks operating on\ndifferent distributions that might be suboptimal for performance. The proposed\nsolution does not require saliency masks, instead it uses a regularization loss\nto prevent the attention maps from spreading to less discriminative object\nregions. Our experiments on the Pascal VOC dataset demonstrate that our\nadversarial approach increases segmentation performance by 2.1 mIoU compared to\nour baseline and by 1.0 mIoU compared to previous adversarial erasing\napproaches.\n
Paper
References (55)
Scroll for more · 38 remaining