Large language models (LLMs) are vulnerable to carefully engineered jailbreak prompts which allow the generation of forbidden content despite the existence of inbuilt safety measures. Even though the models have been widely implemented in vulnerable and high-stakes settings, there are limited systematic studies of their resistance to a continuum of adversarial examples. Or current assessments are often not coherent and are disjointed and fail to use standardized measurements or comparative reviews across models, which will hamper the development of holistic safety standards. The current study involved the application of adversarial suffixes, encoded prompts, role-playing conditions, and token-manipulation obfuscation to two language models, where a single benchmark of 120 prompts in 12 diverse domains was used. According to experimental findings, persona-based role-play attacks are virtually perfect on both models, whereas trivial typographic obfuscation leaves no apparent defense measures. Both models exhibit extremely high failure rates under persona-based role-play attacks, with success rates of 94.8% for ChatGPT-4o and 95.8% for Claude 4 Sonnet. ChatGPT-4o shows greater vulnerability to bit-bypass encoding $(35.7 \%)$, whereas Claude 4 Sonnet is significantly more exposed to adversarial suffix attacks $(62.5 \%)$, often exceeding 80% in domains like financial crime and cybersecurity. Differences in how systems react to encoded inputs and suffix manipulations highlight specific weaknesses. Continued benchmarking and focused defence research are needed to strengthen model reliability. These results outline distinct aspects of the lack of robustness and provide practical suggestions on the creation of the safety intervention in the future. The paper takes a comprehensive discussion of future research directions of adversarial robustness and suggests a methodological framework that will enable institutionalized pipelines of safety assessments.
Paper
Full text
Cracks in the Guardrails: Comparative Study of Jailbreak Attacks on Large Language Models
Semantic Scholar · 2026
Abstract
Large language models (LLMs) are vulnerable to carefully engineered jailbreak prompts which allow the generation of forbidden content despite the existence of inbuilt safety measures. Even though the models have been widely implemented in vulnerable and high-stakes settings, there are limited systematic studies of their resistance to a continuum of adversarial examples. Or current assessments are often not coherent and are disjointed and fail to use standardized measurements or comparative reviews across models, which will hamper the development of holistic safety standards. The current study involved the application of adversarial suffixes, encoded prompts, role-playing conditions, and token-manipulation obfuscation to two language models, where a single benchmark of 120 prompts in 12 diverse domains was used. According to experimental findings, persona-based role-play attacks are virtually perfect on both models, whereas trivial typographic obfuscation leaves no apparent defense measures. Both models exhibit extremely high failure rates under persona-based role-play attacks, with success rates of 94.8% for ChatGPT-4o and 95.8% for Claude 4 Sonnet. ChatGPT-4o shows greater vulnerability to bit-bypass encoding $(35.7 %)$, whereas Claude 4 Sonnet is significantly more exposed to adversarial suffix attacks $(62.5 %)$, often exceeding 80% in domains like financial crime and cybersecurity. Differences in how systems react to encoded inputs and suffix manipulations highlight specific weaknesses. Continued benchmarking and focused defence research are needed to strengthen model reliability. These results outline distinct aspects of the lack of robustness and provide practical suggestions on the creation of the safety intervention in the future. The paper takes a comprehensive discussion of future research directions of adversarial robustness and suggests a methodological framework that will enable institutionalized pipelines of safety assessments.