Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory

Despite existing security alignment mechanisms, Large Language Models (LLMs) remain vulnerable to jailbreak attacks under static defenses. To address this, we propose a novel jailbreak attack and defense optimization framework based on Red-Blue Team dynamic game theory. This framework establishes an automated adversarial mechanism between red and blue teams, achieving closed-loop optimization of instruction generation, semantic interception, and strategy evolution. The red team employs reinforcement learning and semantic enhancement techniques to generate highly diverse and concealed adversarial instructions; the blue team constructs a three-tier defense system—"Rule Filtering-Deep Learning-Dynamic Updating"—to enable real-time interception of harmful instructions and continuous model reinforcement. Experimental results demonstrate that our framework achieves significant attack effectiveness across various models including Vicuna-7B, Llama-2-7B, GPT-3.5-turbo, and GPT-4, with a peak Attack Success Rate (ASR) of 98.7%. On the defense side, it reduces the ASR of GCG attacks to 0.9%, validating the efficacy of the dynamic game mechanism in enhancing model robustness.

Paper

Full text

PDF

Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory

Semantic Scholar · 2026

Abstract

Despite existing security alignment mechanisms, Large Language Models (LLMs) remain vulnerable to jailbreak attacks under static defenses. To address this, we propose a novel jailbreak attack and defense optimization framework based on Red-Blue Team dynamic game theory. This framework establishes an automated adversarial mechanism between red and blue teams, achieving closed-loop optimization of instruction generation, semantic interception, and strategy evolution. The red team employs reinforcement learning and semantic enhancement techniques to generate highly diverse and concealed adversarial instructions; the blue team constructs a three-tier defense system—"Rule Filtering-Deep Learning-Dynamic Updating"—to enable real-time interception of harmful instructions and continuous model reinforcement. Experimental results demonstrate that our framework achieves significant attack effectiveness across various models including Vicuna-7B, Llama-2-7B, GPT-3.5-turbo, and GPT-4, with a peak Attack Success Rate (ASR) of 98.7%. On the defense side, it reduces the ASR of GCG attacks to 0.9%, validating the efficacy of the dynamic game mechanism in enhancing model robustness.

Similar papers

© 2026 NYSGPT2525 LLC