Recent advances in self-supervised learning (SSL) have largely closed the gap\nwith supervised ImageNet pretraining. Despite their success these methods have\nbeen primarily applied to unlabeled ImageNet images, and show marginal gains\nwhen trained on larger sets of uncurated images. We hypothesize that current\nSSL methods perform best on iconic images, and struggle on complex scene images\nwith many objects. Analyzing contrastive SSL methods shows that they have poor\nvisual grounding and receive poor supervisory signal when trained on scene\nimages. We propose Contrastive Attention-Supervised Tuning(CAST) to overcome\nthese limitations. CAST uses unsupervised saliency maps to intelligently sample\ncrops, and to provide grounding supervision via a Grad-CAM attention loss.\nExperiments on COCO show that CAST significantly improves the features learned\nby SSL methods on scene images, and further experiments show that CAST-trained\nmodels are more robust to changes in backgrounds.\n
Paper
References (66)
Scroll for more · 38 remaining