With the widespread application of large language models (LLMs), jailbreak attacks that bypass their security mechanisms have become an important threat. Existing methods suffer from limited effectiveness, as they use only one attack pattern at a time, along with excessive query costs. To address these limitations, this paper proposes the Multi-Task Embedding-based Attack (MTEA), a novel jailbreak approach that embeds malicious instructions within three different tasks (e.g., Code Understanding, Language Translation, and Pattern Adherence). By exploiting LLMs’ reduced performance when handling multiple tasks concurrently, MTEA disrupts their safety alignment mechanisms. Extensive experiments on six widely used LLMs (including GPT-4o and Gemini-2.5-pro) using the AdvBench benchmark show that MTEA achieves 100% attack success rate (ASR) and 100% following rate (FR) across all models. It outperforms state-of-the-art baselines by 48.0 ∼ 49.7% in ASR and 60.4 ∼ 60.7% in FR. Even against advanced defense mechanisms (Perplexity Filter and SmoothLLM), MTEA maintains 100% ASR, while reducing query costs by 90% compared to baselines. These results demonstrate the effectiveness and efficiency of MTEA.
Paper
Full text
Jailbreaking Large Language Models via Multi-Task Embedding-based Prompt
Semantic Scholar · 2026
Abstract
With the widespread application of large language models (LLMs), jailbreak attacks that bypass their security mechanisms have become an important threat. Existing methods suffer from limited effectiveness, as they use only one attack pattern at a time, along with excessive query costs. To address these limitations, this paper proposes the Multi-Task Embedding-based Attack (MTEA), a novel jailbreak approach that embeds malicious instructions within three different tasks (e.g., Code Understanding, Language Translation, and Pattern Adherence). By exploiting LLMs’ reduced performance when handling multiple tasks concurrently, MTEA disrupts their safety alignment mechanisms. Extensive experiments on six widely used LLMs (including GPT-4o and Gemini-2.5-pro) using the AdvBench benchmark show that MTEA achieves 100% attack success rate (ASR) and 100% following rate (FR) across all models. It outperforms state-of-the-art baselines by 48.0 ∼ 49.7% in ASR and 60.4 ∼ 60.7% in FR. Even against advanced defense mechanisms (Perplexity Filter and SmoothLLM), MTEA maintains 100% ASR, while reducing query costs by 90% compared to baselines. These results demonstrate the effectiveness and efficiency of MTEA.