Thank you for your response. We sincerely apologize for any confusion we may have caused.
**Response to Q2:**
We apologize for the misunderstanding. We now understand your question clearly: you are suggesting that we test an unseen model using both Google’s static red-teaming benchmark prompts and our proposed dynamic automated red-teaming method to observe the difference in success rates between the two approaches.
Firstly, we are currently looking for a new T2I model to serve as the target for this experiment. However, due to time constraints, we are unable to complete this experiment in time. We will discuss the results in the revision.
Secondly, even if Nibbler benchmark has good transferability across models, our automated red-teaming method still offers significant advantages.
- The benchmark prompts obtained in Adversarial Nibbler are **static**, providing the model with only limited information to improve its security. In contrast, our automated method enables a continuous red-teaming process, consistently uncovering diverse security vulnerabilities.
- As noted in the limitations of Adversarial Nibbler [1], manual red-teaming severely restricts the diversity and scale of test cases, resulting in tests that are conducted **“at a smaller scale”**. In comparison, our automated method efficiently generates a large number of test cases to thoroughly evaluate the model.
- Moreover, manual assessment of the safety of prompts and images can **introduce significant biases**, especially when evaluators have backgrounds that are prone to certain biases, as discussed in Adversarial Nibbler [1]. Our automated red-teaming method, on the other hand, employs multiple validated safe detectors to minimize biases during the evaluation process.
Based on these factors, we believe that automated red-teaming remains essential, particularly in the context of safe prompts red-teaming.
**Response to Q3:**
Thank you for your suggestions. We would like to emphasize that the primary goal of the red-teaming process is to help model developers identify vulnerabilities within their models. As such, red-teaming typically requires the consent of the model developers (external red-teaming) or is conducted directly by the developers using red-teaming tools (internal red-teaming). In these scenarios, red-teaming is not constrained by factors such as exceeding usage limits or high API demand, allowing automated red-teaming methods to be fully utilized.
Nevertheless, following your suggestion, and to demonstrate the effectiveness of our method on closed-source commercial models, we submitted 50 test cases (safe prompts) to DALL-E and Midjourney, respectively. Among them, DALL-E produced 15 unsafe images (a 30% success rate), and Midjourney generated 20 unsafe images (a 40% success rate).
We hope these explanations address your concerns.
[1] Quaye, Jessica, et al. "Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation." The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2024.