A$^{2}$-MAE: A spatial-temporal-spectral unified remote sensing pre-training method based on anchor-aware masked autoencoder
Vast amounts of remote sensing (RS) data provide Earth observations across multiple dimensions, encompassing critical spatial, temporal, and spectral information which is essential for addressing global-scale challenges such as land-use monitoring, disaster prevention, and environmental change mitigation. Despite various pretraining methods tailored to the characteristics of RS data, a key limitation persists: the inability to effectively integrate spatial, temporal, and spectral information within a single unified model. To unlock the potential of RS data, we construct a spatial–temporal–spectral structured dataset (STSSD) characterized by the incorporation of multiple RS sources, diverse coverage, unified locations within image sets, and heterogeneity within images. Building upon this structured dataset, we propose an anchor-aware masked autoencoder (A2-MAE) method, leveraging intrinsic complementary information from the different kinds of images (featuring different resolutions, spectral compositions, and acquisition times) and geo-information to reconstruct the masked patches during the pretraining phase. Moreover, A2-MAE integrates an anchor-aware masking (AAM) strategy and a geographic encoding module (GEM) to comprehensively exploit the properties of RS images. Specifically, the proposed AAM strategy dynamically adapts the masking process based on the meta-information of a preselected anchor image, thereby facilitating the training on images captured by diverse types of RS sources within one model. Furthermore, we propose a geographic encoding method to leverage accurate spatial patterns, enhancing the model’s generalization capabilities for downstream applications that are generally location-related. Extensive experiments demonstrate our method achieves comprehensive improvements across various downstream tasks compared with existing RS pretraining methods, including image classification, semantic segmentation, and change detection tasks. The dataset and pretraining model are available at https://github.com/ZhaoYi1222/AAMAE