Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
Large Language Models (LLMs) remain vulnerable to adversarial prompts that elicit harmful responses. While many existing safety systems are better able to detect overtly nonsensical attack strings, human-readable prompts embedded in plausible situational contexts remain harder to identify and evaluate. This paper presents an empirical investigation of human-readable, situation-driven adversarial prompts for assessing LLM robustness. First, we use movie scripts as situational contexts (e.g., crime narratives) to construct natural-looking prompts that bypass safety mechanisms. Second, we transform adversarial gibberish into coherent, innocuous-appearing text that retains exploitation capability within these contextual frameworks. Third, we enhance the AdvPrompter framework with p-nucleus sampling to generate diverse human-readable attacks, substantially improving success rates against models including GPT-3.5 and Gemma-7b. We validate our approach through multi-method evaluation: automated harmfulness scoring (GPT-4o-mini), independent human assessment by 10 raters across 80 samples, and Elo rating meta-analysis of judge reliability. These findings highlight the need for safety mechanisms that can detect not only nonsensical jailbreak strings, but also coherent adversarial content embedded in realistic narrative contexts. The code for this study is publicly available.1
Paper
References (58)
Scroll for more · 38 remaining