Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models

Agentic methods have emerged as a powerful paradigm that enhances reasoning, collaboration, and adaptive control, enabling systems to coordinate and solve complex tasks. We extend this paradigm to safety alignment by introducing Agentic Moderation, a model-agnostic framework that leverages specialized agents to defend multimodality models against jailbreak attacks. Unlike prior moderator that operate as a static layer over inputs or outputs and provide binary classifications (safe or unsafe), our method integrates cooperative agents, including Shield, Responder, Evaluator, and Reflector, to achieve context-aware and interpretable moderation. Extensive experiments across five datasets and four representative large vision-language models (LVLMs) demonstrate that our approach reduces the Attack Success Rate (ASR) by 7-19%, maintains a stable Non-Following Rate (NF), and improves the Refusal Rate (RR) by 4-20%, achieving robust and well-balanced safety performance. By harnessing the flexibility and reasoning capacity of agentic architectures, Agentic Moderation provides modular, scalable, and fine-grained safety enforcement, highlighting the potential of agentic systems as a foundation for automated safety governance. This works Code is available at [https://github.com/adaren100/moderator

Paper

References (32)

Scroll for more · 20 remaining

Similar papers

© 2026 NYSGPT2525 LLC