Diffusion models have emerged as the state-of-the-art generative paradigm in image synthesis. However, their powerful generative and representational capabilities also make them highly susceptible to backdoor attacks, where specific triggers embedded in the input can manipulate the model to produce attacker-specified outputs. Traditional triggers typically rely on explicit perturbations in low-dimensional space, such as patch overlays, which are relatively easy to detect and defend against. To explore more covert and effective backdoor injection strategies, we propose a novel method, Bad-Pose Diffusion (Bad-PoseDiff), which uses pose features in input images as triggers. Since pose information is a high-dimensional semantic feature that manifests in diverse and non-fixed patterns, it is difficult for models to recognize directly. We therefore introduce the ControlNet module to inject pose-guided conditioning signals into the diffusion model during training. To the best of our knowledge, this is the first work to incorporate pose-based triggers into generative diffusion models via ControlNet. Notably, even when ControlNet is detached during inference, the backdoored model continues to recognize and respond to pose triggers, indicating that the backdoor has been deeply implanted into the backbone of the diffusion model. Experimental results show that Bad-PoseDiff effectively evades all existing defense mechanisms, achieving a 0% Backdoor Detection Rate (BDR) across all evaluated frameworks, while preserving high-quality outputs.
Paper
Full text
Bad-PoseDiff: Pose-Guided Backdoor Triggering in Diffusion Models
OpenAlex · Generative Adversarial Networks and Image Synthesis · 2025
Abstract
Diffusion models have emerged as the state-of-the-art generative paradigm in image synthesis. However, their powerful generative and representational capabilities also make them highly susceptible to backdoor attacks, where specific triggers embedded in the input can manipulate the model to produce attacker-specified outputs. Traditional triggers typically rely on explicit perturbations in low-dimensional space, such as patch overlays, which are relatively easy to detect and defend against. To explore more covert and effective backdoor injection strategies, we propose a novel method, Bad-Pose Diffusion (Bad-PoseDiff), which uses pose features in input images as triggers. Since pose information is a high-dimensional semantic feature that manifests in diverse and non-fixed patterns, it is difficult for models to recognize directly. We therefore introduce the ControlNet module to inject pose-guided conditioning signals into the diffusion model during training. To the best of our knowledge, this is the first work to incorporate pose-based triggers into generative diffusion models via ControlNet. Notably, even when ControlNet is detached during inference, the backdoored model continues to recognize and respond to pose triggers, indicating that the backdoor has been deeply implanted into the backbone of the diffusion model. Experimental results show that Bad-PoseDiff effectively evades all existing defense mechanisms, achieving a 0% Backdoor Detection Rate (BDR) across all evaluated frameworks, while preserving high-quality outputs.