MAJD: Intent-Aware Multi-Agent Framework for Jailbreak Defense

Large Language Models (LLMs) have shown remarkable capabilities across diverse applications, enabling broad user interaction. Despite advances in AI alignment, their susceptibility to adversarial jailbreak attacks presents serious safety challenges. To mitigate these risks, we propose MAJD, a Multi-Agent Jailbreak Defense Framework that operates entirely at inference time without retraining. Each agent performs specialized tasks: the Coordinator routes inputs based on semantic ambiguity, the Mutation Agent compresses prompts with predefined rules, the Analyzer conducts intent analysis and response generation, and the Revalidation Agent performs secondary response classification. By incorporating rulebased mutation and intent-based reasoning, our framework identifies the core malicious intent in both single-turn and multi-turn jailbreak attempts. Experimental results across multiple attack methods (AutoDAN, GPTFuzzer, PAIR, and MHJ) and across multiple target models demonstrate that MAJD significantly improves detection success rates. These results validate the efficiency and generalizability of MAJD as a scalable jailbreak defense solution for LLM applications.

Paper

Full text

PDF

MAJD: Intent-Aware Multi-Agent Framework for Jailbreak Defense

Semantic Scholar · Computer Science · 2025

Abstract

Large Language Models (LLMs) have shown remarkable capabilities across diverse applications, enabling broad user interaction. Despite advances in AI alignment, their susceptibility to adversarial jailbreak attacks presents serious safety challenges. To mitigate these risks, we propose MAJD, a Multi-Agent Jailbreak Defense Framework that operates entirely at inference time without retraining. Each agent performs specialized tasks: the Coordinator routes inputs based on semantic ambiguity, the Mutation Agent compresses prompts with predefined rules, the Analyzer conducts intent analysis and response generation, and the Revalidation Agent performs secondary response classification. By incorporating rulebased mutation and intent-based reasoning, our framework identifies the core malicious intent in both single-turn and multi-turn jailbreak attempts. Experimental results across multiple attack methods (AutoDAN, GPTFuzzer, PAIR, and MHJ) and across multiple target models demonstrate that MAJD significantly improves detection success rates. These results validate the efficiency and generalizability of MAJD as a scalable jailbreak defense solution for LLM applications.

Similar papers

© 2026 NYSGPT2525 LLC