Thanks for your response.
> If you have some thoughts on why this _(standard datasets are hard to find)_ might be the case, please do share. Might it be the case that collecting such data is not easy / costly?
Yes, standard datasets are hard to find partly because their curation is difficult when using a conventional annotation pipeline. However, with increasing interest in learning from explanations as a method for training reliable models [3, 4, 6, 7, 8], we are witnessing growing number of relevant datasets and techniques for efficient data curation. We answered more elaborately on advances in curation of relevant datasets in our response to Reviewer k7a7.
Their question and our response to their question are relevant here, which we are pasting here for easy reference.
----
Reviewer k7a7:
> Obtaining human-specified masks is at best a lot of effort …
Our response:
We agree that manually specifying explanation masks can be impractical. However, the procedure can be automated if the nuisance/irrelevant feature occurs systematically or if it is easy to recognize, which may then be obtained automatically using a procedure similar to [1,2]. A recent effort called Salient-Imagenet used neuron activation maps to scale curation of such human-specified masks to Imagenet-scale [3, 4]. These efforts may be seen as a proof-of-concept for obtaining richer annotations beyond content labels, and towards better defined tasks.
-----
We hope this answers your question. We will make sure to include these details in the main paper. We wish to also highlight that we evaluated a new dataset (Salient-Imagenet that was shared in the general response), which was inspired by your concern on standard datasets. Thanks again for your time and comment.
References.
[1] Liu, Evan Z., et al. "Just train twice: Improving group robustness without training group information." International Conference on Machine Learning. PMLR, 2021.
[2] Rieger, Laura, et al. "Interpretations are useful: penalizing explanations to align neural networks with prior knowledge." International conference on machine learning. PMLR, 2020.
[3] Singla, Sahil, and Soheil Feizi. "Salient ImageNet: How to discover spurious features in Deep Learning?." arXiv preprint arXiv:2110.04301 (2021).
[4] Singla, Sahil, Mazda Moayeri, and Soheil Feizi. "Core risk minimization using salient imagenet." arXiv preprint arXiv:2203.15566 (2022).
[5] Ross, Andrew Slavin, Michael C. Hughes, and Finale Doshi-Velez. "Right for the right reasons: Training differentiable models by constraining their explanations." arXiv preprint arXiv:1703.03717 (2017).
[6] Pukdee, Rattana, et al. "Learning with Explanation Constraints." arXiv preprint arXiv:2303.14496 (2023).
[7] Miao, Kevin, et al. "Prior Knowledge-Guided Attention in Self-Supervised Vision Transformers." arXiv preprint arXiv:2209.03745 (2022).
[8] Yang, Ziyan, et al. "Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.