We thank the reviewer for the comments. Here, we address the points on metadata and LLMs.
**Metadata**
In the absence of ground truth metadata, one can leverage large multimodal models (LMMs), for instance in segmentation or open-vocabulary detection, and generate such concepts. However, there remains the possibility of noisy metadata that may contain unreliable artefacts, in addition to the biases within these models themselves. In this work, we restrict ourselves to available, high-quality, ground truth concepts, to showcase the usefulness of our framework. Pursuing open-vocabulary models to generate metadata is an interesting direction of future work. In general, ConBias is not constrained by how the metadata is obtained, but in the quality of metadata obtained. Future developments in LLMs/LMMs that can generate high quality concept metadata can be seamlessly integrated into the ConBias framework. We will be sure to update the paper with this discussion.
Our core contribution with ConBias is not the metadata stage, which we assume to be available and high-quality, similar to other works in the past [A, B]. Our core contribution is the diagnosis and debiasing of datasets with ConBias, which leads to significant improvements on multiple datasets with respect to the current state-of-the-art. As requested by the reviewer, we have also provided diagnosis results on a more complex dataset such as Imagenet-1k.
[A] Wu, Shirley, et al. "Discover and cure: Concept-aware mitigation of spurious correlation." In ICML, 2023.
[B] Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. Diversify your vision datasets with automatic diffusion-based augmentation. Advances in Neural Information Processing Systems, 36, 2024
**LLMs**
We reiterate that, as mentioned in lines 36-39, relying on LLMs to generate diverse, unbiased descriptions is problematic since LLMs themselves may be biased, and such generation is not controllable. This issue of relying on LLMs has also been addressed in the ALIA paper (Section 6) [A], and it is precisely this issue that we fix with ConBias. By leveraging the concept graph, ConBias can generate a debiased dataset in a controlled and interpretable manner.
[A] Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. Diversify your vision datasets with automatic diffusion-based augmentation. Advances in Neural Information Processing Systems, 36, 2024