Hybrid Transformer Network for Change Detection Under Self-Supervised Pretraining

This paper presents a Siamese network architecture based on a multi-scale hybrid convolution-Transformer (CTUNet) for Change Detection (CD) in a pair of co-registered optical remote sensing images. Different form CD frameworks based on convolution neural networks (CNNs) and pure Transformer networks, this method combines a convolution-Transformer hybrid encoder with a multi-scale change information extraction decoder in a Siamese network architecture. It overcomes the inherent limitations of CNN and Transformer and effectively integrates the multi-scale information required for accurate CD. To learn better discriminative representations from various scales, we propose a masked auto-encoder scheme (CTMAE) to adapt to building targets with varying morphological scales, further unleashing the potential of CTUNet. Experiments on two CD datasets show that the proposed self-supervised pre-trained hybrid convolution-Transformer CTUNet architecture achieves better CD performance than previous methods.

Paper

Full text

PDF

Hybrid Transformer Network for Change Detection Under Self-Supervised Pretraining

Semantic Scholar · Computer Science · 2023

Abstract

This paper presents a Siamese network architecture based on a multi-scale hybrid convolution-Transformer (CTUNet) for Change Detection (CD) in a pair of co-registered optical remote sensing images. Different form CD frameworks based on convolution neural networks (CNNs) and pure Transformer networks, this method combines a convolution-Transformer hybrid encoder with a multi-scale change information extraction decoder in a Siamese network architecture. It overcomes the inherent limitations of CNN and Transformer and effectively integrates the multi-scale information required for accurate CD. To learn better discriminative representations from various scales, we propose a masked auto-encoder scheme (CTMAE) to adapt to building targets with varying morphological scales, further unleashing the potential of CTUNet. Experiments on two CD datasets show that the proposed self-supervised pre-trained hybrid convolution-Transformer CTUNet architecture achieves better CD performance than previous methods.

Similar papers

© 2026 NYSGPT2525 LLC