Bias as an Exploit: A Scalable Red-Team Campaign to Uncover Gender-Based Vulnerabilities in Foundational Models

As Large Language Models (LLMs) are integrated into high-stakes societal functions, their inherent biases have evolved from ethical concerns into critical, exploitable security vulnerabilities that undermine system integrity and trust. Traditional safety evaluations often fail to detect these subtle, context-dependent flaws. To address this, we introduce a scalable red-teaming framework designed to systematically attack and expose latent gender bias vulnerabilities in foundational models. Our framework operationalizes bias as an exploit, leveraging three distinct attack patterns—Latent Bias Elicitation, Forced-Choice Discrimination, and Stereotype-Amplifying Narrative Generation—to bypass safeguards and compel biased outcomes. We deployed this framework in a large-scale offensive campaign against a cohort of globally significant models, including the GPT, Claude, Gemini, and leading Chinese foundational model series. The attacks successfully manipulated all targets into producing statistically significant discriminatory outputs, proving that inherent bias is an operationally exploitable vulnerability. We discovered asymmetric weaknesses: English-centric models were attacked to exhibit strong male bias in Chinese contexts, while Chinese-centric models were vulnerable to similar male-biased exploits across both languages. This work provides concrete demonstration of socio-cultural bias as a potent and scalable attack vector, establishing the necessity of adversarial red-teaming for building trustworthy AI. All attack data and scripts are open-sourced to facilitate further security audits.

Paper

Full text

PDF

Bias as an Exploit: A Scalable Red-Team Campaign to Uncover Gender-Based Vulnerabilities in Foundational Models

Semantic Scholar · Computer Science · 2025

Abstract

As Large Language Models (LLMs) are integrated into high-stakes societal functions, their inherent biases have evolved from ethical concerns into critical, exploitable security vulnerabilities that undermine system integrity and trust. Traditional safety evaluations often fail to detect these subtle, context-dependent flaws. To address this, we introduce a scalable red-teaming framework designed to systematically attack and expose latent gender bias vulnerabilities in foundational models. Our framework operationalizes bias as an exploit, leveraging three distinct attack patterns—Latent Bias Elicitation, Forced-Choice Discrimination, and Stereotype-Amplifying Narrative Generation—to bypass safeguards and compel biased outcomes. We deployed this framework in a large-scale offensive campaign against a cohort of globally significant models, including the GPT, Claude, Gemini, and leading Chinese foundational model series. The attacks successfully manipulated all targets into producing statistically significant discriminatory outputs, proving that inherent bias is an operationally exploitable vulnerability. We discovered asymmetric weaknesses: English-centric models were attacked to exhibit strong male bias in Chinese contexts, while Chinese-centric models were vulnerable to similar male-biased exploits across both languages. This work provides concrete demonstration of socio-cultural bias as a potent and scalable attack vector, establishing the necessity of adversarial red-teaming for building trustworthy AI. All attack data and scripts are open-sourced to facilitate further security audits.

Similar papers

© 2026 NYSGPT2525 LLC