MultiTaskVIF: Segmentation-oriented visible and infrared image fusion via multi-task learning

Visible and infrared image fusion (VIF) has attracted significant attention in recent years. Traditional VIF methods primarily focus on generating fused images with high visual quality, while recent advancements increasingly emphasize incorporating semantic information into the fusion model during training. However, most existing segmentation-oriented VIF methods adopt a cascade structure comprising separate fusion and segmentation models, leading to increased network complexity and redundancy. This raises a critical question: can we design a more concise and efficient structure to integrate semantic information directly into the fusion model during training? Inspired by multi-task learning (MTL), we propose a concise and universal training framework, MultiTaskVIF, for segmentation-oriented VIF models. In this framework, we introduce a multi-task head decoder (MTH) that leverages the segmentation head to inject informative semantics into the fusion branch during training. Unlike previous cascade training frameworks that necessitate joint training with a complete segmentation model, MultiTaskVIF enables the existing image fusion model to learn semantic features by simply replacing its decoder with the proposed MTH. Moreover, by combining a hard parameter sharing MTL architecture with an attention mechanism, our method not only reduces network complexity and redundancy but also enhances the fusion model's ability to extract segmentation-relevant features, thereby improving the quality of the fused image for downstream segmentation tasks. Extensive experimental evaluations validate the effectiveness of the proposed method. Our code will be released upon acceptance.

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC