While diffusion distillation has enabled one-step generation through methods like Diff-Instruct [17] and Variational Score Distillation [34], adapting distilled models to emerging new controls - such as novel structural constraints or latest user preferences - remains challenging. Conventional approaches typically require modifying the base diffusion model and redistilling it - a process that is both computationally intensive and time-consuming. To address these challenges, we introduce Joint Distribution Matching (JDM), a novel approach that minimizes the reverse KL divergence between image-condition joint distributions. By deriving a tractable upper bound, JDM decouples fidelity learning from condition learning. This asymmetric distillation scheme enables our one-step student to handle controls unknown to the teacher model and facilitates improved classifier-free guidance ($C F G$) usage and seamless integration of human feedback learning (HFL). Experimental results demonstrate that JDM surpasses baseline methods such as multi-step ControlNet [45] by mere one-step in most cases, while achieving state-of-the-art performance in one-step text-to-image synthesis through improved usage of $C F G$ or HFL integration.
Paper
References (45)
Scroll for more · 33 remaining