Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling
Earth observation (EO) foundation models have emerged as an effective approach to derive latent representations of the Earth system from various remote sensing sensors. These models produce embeddings that can be used as analysis-ready datasets, enabling the modeling of ecosystem dynamics without extensive sensor-specific preprocessing. However, existing models typically operate over large spatial or temporal contexts, limiting their use for ecological analyses that require modeling across task-specific spatial scales and diverse temporal conditions. To overcome these limitations, we propose a representation learning framework that integrates different EO modalities into a unified feature space on a high-resolution spatio-temporal grid. We introduce the framework using Sentinel-1 and Sentinel-2 data as representative modalities. Our approach produces spatio-temporally explicit embeddings aligned with the native 10-m spatial grid and the temporal frequency of cloud-free Sentinel-2 acquisitions. Each sensor is first modeled independently to capture its sensor-specific characteristics. Their representations are then combined in a shared model. This two-stage design enables modality-specific optimization and easy extension to new sensors, retaining pretrained encoders while training only the added fusion layers. The fusion model achieved a validation reconstruction MAE of 0.0004 for Sentinel-1 and 0.0039 for Sentinel-2 on normalized inputs, indicating accurate cross-modal reconstruction. Qualitative analyses reveal that the learned embeddings exhibit high spatial and semantic consistency across heterogeneous landscapes. Quantitative evaluation in modeling gross primary production reveals that they encode ecologically meaningful patterns and preserve the temporal frequency required for fine-scale environmental analyses, with an NRMSE of 0.0917. Overall, the proposed framework provides a flexible, analysis-ready representation learning approach for environmental applications requiring diverse spatial and temporal contexts.
Paper
References (82)
Scroll for more · 38 remaining