Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling semantic controls from generation pipelines, a fundamental bottleneck lies in their inability to simultaneously capture holistic facial structure and region-level semantics, resulting in coarse, entangled representations that hinder fine-grained attribute control. We revisit this challenge through the lens of expert specialization and introduce Face-MoGLE, a novel framework featuring: (1) Semantic-decoupled latent modeling through mask-conditioned space factorization, enabling precise attribute synthesis; (2) A mixture of global and local experts, motivated by the need to decouple holistic structural reasoning from region-specific semantic refinement, achieving fine-grained controllability that monolithic encoders cannot provide; (3) A diffusion-aware dynamic gating network that produces time- and spatially dependent coefficients evolving with the denoising process. In addition, to alleviate the scarcity of evaluation benchmarks for multimodal face generation, we further extend two existing face datasets into enriched multimodal versions through a semiautomatic annotation pipeline, resulting in MM-FFHQ-Female (derived from FFHQ-Text) and MM-FairFace-HQ (derived from FairFace). These datasets provide both textual attributes and semantic masks, facilitating generalization assessment under di verse stylistic variations. Extensive experiments demonstrate that Face-MoGLE significantly outperforms state-of-the-art (SOTA) controllable face generation methods in both unimodal and multimodal settings, achieving superior image quality, semantic alignment, aesthetic preference, and zero-shot generalization. Beyond conventional metrics, we additionally evaluate Face MoGLE using SOTA deepfake detectors to confirm that Face MoGLE produces highly realistic images capable of confusing forgery detection models, underscoring both the strength of our method and the importance of responsible AI considerations. We release the code, pretrained models, and related data at https://github.com/XavierJiezou/Face-MoGLE.
Paper
References (60)
Scroll for more · 38 remaining