In-sample Contrastive Learning and Consistent Attention for Weakly Supervised Object Localization

Weakly supervised object localization (WSOL) aims to localize the target\nobject using only the image-level supervision. Recent methods encourage the\nmodel to activate feature maps over the entire object by dropping the most\ndiscriminative parts. However, they are likely to induce excessive extension to\nthe backgrounds which leads to over-estimated localization. In this paper, we\nconsider the background as an important cue that guides the feature activation\nto cover the sophisticated object region and propose contrastive attention\nloss. The loss promotes similarity between foreground and its dropped version,\nand, dissimilarity between the dropped version and background. Furthermore, we\npropose foreground consistency loss that penalizes earlier layers producing\nnoisy attention regarding the later layer as a reference to provide them with a\nsense of backgroundness. It guides the early layers to activate on objects\nrather than locally distinctive backgrounds so that their attentions to be\nsimilar to the later layer. For better optimizing the above losses, we use the\nnon-local attention blocks to replace channel-pooled attention leading to\nenhanced attention maps considering the spatial similarity. Last but not least,\nwe propose to drop background regions in addition to the most discriminative\nregion. Our method achieves state-of-theart performance on CUB-200-2011 and\nImageNet benchmark datasets regarding top-1 localization accuracy and\nMaxBoxAccV2, and we provide detailed analysis on our individual components. The\ncode will be publicly available online for reproducibility.\n

Paper

References (32)

Scroll for more · 20 remaining

Similar papers

© 2026 NYSGPT2525 LLC