We present TerraMind, the first any-to-any generative, multimodal deep learning model for Earth observation (EO). Unlike other approaches, TerraMind is pretrained on dualscale representations combining both token-level and pixellevel data across modalities. On a token level, TerraMind encodes high-level contextual information to learn crossmodal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. In this paper, we demonstrate that (i) TerraMind achieves beyond state-of-the-art performance in communitystandard benchmarks, (ii) TerraMind can leverage “thinking in modalities” (TiM)-the capability of generating additional artificial data during finetuning and inference to improve the model output-and (iii) TerraMind's dual-scale early fusion approach results in well-structured embedding spaces. Models and code have been open-sourced at https://huggingface.co/ibm-esa-geospatial and https://github.com/ibm/terramind.
Paper
References (75)
Scroll for more · 38 remaining