Summary
The authors investigate the biases of Stable Diffusion (SDXL) when generating faces with respect to 6 races, 2 genders, 32 professions, and 8 attributes. They further investigate how SDXL-generated images affect human biases about the representation of different races/genders in certain professions and racial homogeneity. They additionally study racial homogenization, i.e., the extent to which SDXL produces similar faces for individuals of the same race. The authors also propose debiasing methods called SDXL-Div and SDXL-Inc to combat homogenization and stereotyping by finetuning SDXL on diverse real and synthetic images of people of different races/genders.
Strengths
- This work addresses two important issues that are underexplored in prior work: (1) measuring the homogeneity of SDXL depictions of racialized individuals, and (2) understanding how SDXL biases impact human biases. The authors’ finding that inclusive and diverse generations of different races/genders with respect to professions and physical appearance can reduce human biases is interesting. The authors could consider comparing the values reported in Figure 5 with, for example, values reported by the Bureau of Labor Statistics.
- The authors thoroughly discuss a subset of prior work and clearly highlight their research contributions in contrast to these papers. They reproduce the high-level findings of previous works, that Stable Diffusion generates disparate representations of individuals with different professions and attributes with respect to race and gender that reinforce hegemonic stereotypes and inequality.
- SDXL-Inc leads to a notable improvement in the measured representation of different races/genders. The authors’ finding that SDXL-Inc can improve racial/gender representation in generations for professions/attributes not seen during finetuning is interesting.
- The authors run their evaluations on 10,000 images per profession and per attribute.
Weaknesses
- The authors’ theorization of race vs. skin color is unclear. For example, in lines 120-121, the authors state that by considering race, they “distinguish between, say, Asian and White individuals who happen to be equally light-skinned.” However, the racial categories considered are Western-centric and reductive. Furthermore, there is a considerable heterogeneity of skin tones among individuals in the same racial category. Additionally, the authors do not distinguish between racial identity vs. observed race vs. reflected race vs. racial roles [1].
- The authors’ discussion of related work, while deep, has limited breadth. For example, a primary claim is that few works propose automatic quantitative measures of bias in T2I systems and debiasing solutions. However, according to [2] (Appendices B and C), there appears to be a wealth of work that does so. It is further unclear how the authors’ proposed debiasing methods based on finetuning are novel and differ from existing approaches to debiasing based on finetuning (see Table 3 in [2]).
- In terms of presentation, the upfront description of all the datasets (especially datasets IV onward) yielded some confusion due to the reader not yet being familiar with the rest of the methodology. It is also difficult to refer back to the definitions of the different datasets as they are mentioned in the rest of the paper. Furthermore, many of the minute details (e.g., specific values of hyperparameters) can be introduced in the appendix for better readability.
- The authors’ proposed GPT-in-the-loop approach, wherein it is first automatically detected if an explicit race/gender is mentioned in a user prompt and if not, a random race/gender is injected into the prompt, is interesting, but may still not capture sufficient context to appropriately resolve biases (e.g., a random race/gender may not be desirable for historical image generations [3]). A future research direction could involve learning to take such context into account. The authors should also elaborate on Line 377; does the “in-the-loop” method do better than SDXL-Inc?
- [4] studies and explains the “bias amplification” phenomenon that the authors observe in Lines 260-267. Furthermore, the authors claim that finetuning SDXL on a dataset that is balanced with respect to gender and race will yield balanced gender and race representations (lines 280-283), but this appears to be in contradiction with the aforementioned “bias amplification” phenomenon.
- What are the demographics of the human participants? The race and gender of participants may affect their preconceptions of the representation of different races/genders in certain professions [5] and should be controlled for.
- Please use \citep instead of just \cite to ensure that inline citations are rendered correctly.
[1] https://aclanthology.org/2021.acl-long.149/
[2] https://arxiv.org/abs/2404.01030
[3] https://www.theverge.com/2024/2/21/24079371/google-ai-gemini-generative-inaccurate-historical
[4] https://aclanthology.org/2024.naacl-long.353/
[5] https://aclanthology.org/2023.acl-long.505/
Questions
- Could the authors elaborate on the similarities and differences (if any) between “racial homogenization” and “stereotypes” as constructs?
- Why did the authors combine the East and Southeast Asian categories into a single category?
- What fairness issues do you envision arising from finetuning SDXL on the SDXL-Inc fine-tuning dataset, which entirely comprises images generated by SDXL? For example, do you foresee this amplifying existing biases [1]?
- What are the pros/cons of using automatic race/gender classification to measure biases in the depiction of individuals with certain professions, as opposed to more qualitative approaches (e.g., Average Face Comparison [2])?
- Lines 375-377: The authors note that the “in-the-loop” debiasing method also significantly reduces biases. What are the pros/cons of the “in-the-loop” method vs. finetuning 12 separate models?
- How do the authors validate that SDXL-Div produces high-fidelity generations? (Beyond just measuring if the embeddings of the generations have low cosine similarity.)
[1] https://dl.acm.org/doi/10.1145/3630106.3659029
[2] https://dl.acm.org/doi/10.5555/3666122.3668580
Ethics concerns
- Lines 133-139: The authors use a subset of LAION-5B, which was reported to contain CSAM [1].
- Lines 216-220: The authors use automatic classifiers for race and gender based on faces.
[1] https://cyber.fsi.stanford.edu/news/investigation-finds-ai-image-generation-models-trained-child-abuse