Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

Image fusion aims to blend complementary information from diverse sensing modalities, yet most current methods lack robustness in complex fusion scenarios and cannot flexibly accommodate user intent. We present DiTFuse, the first Diffusion-Transformer (DiT) framework for instruction-driven, dynamic fusion control. Guided by natural-language instructions, DiTFuse flexibly blends multimodal content to enable hierarchical and fine-grained control over fusion dynamics. The training phase employs a multi-degrade-mask-image-modeling (M3) strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ideal reference images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion—as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multiimage fusion scenarios, including instruction-conditioned segmentation.

Paper

References (79)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC