CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders

A vital and rapidly growing application, remote sensing offers vast yet sparsely labeled, spatially aligned multimodal data; this makes self-supervised learning algorithms invaluable. We present CROMA: a framework that combines contrastive and reconstruction self-supervised objectives to learn rich unimodal and multimodal representations. Our method separately encodes masked-out multispectral optical and synthetic aperture radar samples -- aligned in space and time -- and performs cross-modal contrastive learning. Another encoder fuses these sensors, producing joint multimodal encodings that are used to predict the masked patches via a lightweight decoder. We show that these objectives are complementary when leveraged on spatially aligned multimodal data. We also introduce X- and 2D-ALiBi, which spatially biases our cross- and self-attention matrices. These strategies improve representations and allow our models to effectively extrapolate to images up to 17.6x larger at test-time. CROMA outperforms the current SoTA multispectral model, evaluated on: four classification benchmarks -- finetuning (avg. 1.8%), linear (avg. 2.4%) and nonlinear (avg. 1.4%) probing, kNN classification (avg. 3.5%), and K-means clustering (avg. 8.4%); and three segmentation benchmarks (avg. 6.4%). CROMA's rich, optionally multimodal representations can be widely leveraged across remote sensing applications.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer FExC7/10 · confidence 5/52023-06-29

Summary

The paper presents a new SSL representation learning framework remote sensing and earth observation data. The presented framework combines a contrastive objective with a reconstruction objective working on single or multi-modal inputs i.e. multispectral satellite data and synthetic aperture radar data. Cross-modal learning is done by cross-attention of both individually encoded modalities, which are fused and decoded by one lightweight decoder. In a wide range of experiments, the authors are able to demonstrate the proposed approach capability to outperform baseline and current SOTA approaches. Although this work displays rather a novel combination of already known approaches, I think it is interesting given the insightful adaptation of these approaches to the remote sensing and earth observation domain. I really enjoyed reading it. What I really like is the idea to use RPE as presented by the extension of ALiBi towards multispectral 2dim signals allowing to deal with different resolutions of satellite data. This particular characteristic of satellite data is very often neglected.

Strengths

- (S1) The use of RPE with its 2d extension of ALiBi including X-ALiBi. I think this is interesting since it aims to tackle the multi-resolution nature of individual bands of remote sensing data. - (S2) multi-modal representation is optional i.e. it performs well with only one modality if needed. This is in particular interesting for remote sensing scenarios such as natural disasters, where fast response is important but one of the two satellites is not available but will need a couple of days to fly over the target region. - (S3) The wide range of experiments including multiple datasets and downstream tasks all able to demonstrate the outperformance of the presented approach. - (S4) A broad set of ablation studies providing insights about the inter-working and contribution of each component of the proposed method.

Weaknesses

- (W1) see (L1) under limitations.

Questions

- (Q1) How would you extend to more than two modalities? I am asking since in remote sensing and earth observation there are very often multiple modalities / sensors available. How would you model the cross-attention?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

- (L1) The paper does not show the presented approach being capable of generalizing to other problem domains beyond remote sensing or earth observation. I could imagine that there exist other problem domains where multiple sensors are available. Showing how this approach performs in such a scenario would strengthen this work.

Reviewer vsHk3/10 · confidence 5/52023-07-01

Summary

This paper presents a CROMA, a framework that combines contrastive and reconstruction self-supervised objectives to learn rich unimodal and multi-modal representations. CROMA separately encodes masked-out multispectral optical and synthetic aperture radar samples and performs cross-modal contrastive learning. X- and 2D-ALiBi are also introduced to ensure the performance, which spatially biases the cross and self-attention matrices.

Strengths

CROMA aims to address the multi-model learning problem in the remote sensing (RS) community, which is an important and hot topic. Also, many advanced techniques are adopted and combined properly to ensure the final results for different tasks. In sum, the proposed method is feasible.

Weaknesses

However, due to the poor statements and organization, the main ideas of this work are hard to follow. Also, as mentioned above, CROMA combines some existing techniques to deal with its tasks. Thus, its novelty is limited for NIPS. Some detailed comments can be found in “Questions.”

Questions

1. The main contributions are not clear. Many multi-model learning models have been proposed, what are your advantages and own features compared with them? 2. What is FFT? I cannot find its full name in this manuscript. 3. There are three encoders in CROMA. What are the relations between them? What is the rationale behind the freed settings and parameters? How do you decide the input patch sizes? 4. Why do you extend ALiBi to a 2D version? Please explain the necessity. 5. How do you decide the MASK value (i.e., 75%)?

Rating

3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

2 fair

Presentation

1 poor

Contribution

1 poor

Limitations

1. Introduction is chaotic. There is no clear, logical flow between paragraphs, resulting in a disjointed content presentation. 2. A meaningful literature review is missing. The authors only display many published literature. However, the relations between them and CROMA are not clear. Also, the inner relationships of the reviewed literature are confusing. 3. The experimental settings are unclear, preventing the readers from simulating your method. 4. The compared models are not enough, limiting the reliability of the results. 5. The experimental results are discussed as shallow.

Reviewer too94/10 · confidence 5/52023-07-04

Summary

This paper proposes CROMA to align optical and SAR modal images via contrastive learning and reconstruction. Comprehensive experiments on three datasets have demonstrated the effectiveness of CROMA.

Strengths

1. This paper introduces a multi-modal representation using contrastive learning and reconstruction. 2. The proposed CROMA has exceeded the SatMAE, which only used unimodal images. 3. CROMA is more faster and effective than SatMAE.

Weaknesses

1. CROMA only constructs pos-neg samples from different modal images. Why are image patches in different regions of the same modality not used as negative samples? 2. There miss lots of important details. For example, the number of positive and negative samples is not discussed. 3. How is the sampling ratio of positive and negative samples affected? What is the relationship between positive and negative sample sampling and effective inputs in reconstruction? 4. I am worried about the theoretical innovation of the paper. The paper mainly focuses on the application of contrastive learning loss and reconstruction loss to remote sensing multimodal modeling, which is also very common in medical multimodal and multitemporal. As an extension of SatMAE, I don't know whether the innovation of the paper is enough for NeurIPS. Because there already exists many related papers for remote sensing[1][2]. [1] Ayush K, Uzkent B, Meng C, et al. Geography-aware self-supervised learning[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021: 10181-10190. [2] Manas O, Lacoste A, Giró-i-Nieto X, et al. Seasonal contrast: Unsupervised pre-training from uncurated remote sensing data[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021: 9414-9423.

Questions

N/A

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

N/A

Reviewer too92023-08-17

Thanks for the rebuttal

Thank the authors for their rebuttal. It partly addressed my concerns, except for the novelty. This is a good practice to combine contrastive learning and masked image modeling on multi-modal satellite images; however, I cannot regard it as an innovative approach. There is little knowledge improvement for me. Authors continually emphasize that their approach goes beyond SatMAE, which is unfair due to different pre-training data and improper baseline, i.e., SSL4EO v.s. fMoW-Sentinel. The pre-training data should be aligned. The right baseline should be SatMAE+ contrastive learning to illustrate your combination is non-trivial; otherwise, it cannot convince most readers in NeurlPS. I appreciate the extensive empirical results from this manuscript. However, I cannot recommend this manuscript be accepted by NeurlPS at this round, given its limited novelty, unfair comparison, and insufficient theoretical support.

Authorsrebuttal2023-08-17

Thank you for your continued engagement in our work. Regarding novelty, our method is a novel combination of existing methods: cross-modal contrastive learning, multimodal masked autoencoding, and attention with linear biases (ALiBi). Our method was motivated by the intuition that we describe in our paper: (i) that contrastive and reconstructive pretraining objectives learn different representations that might be complementary when combined and (ii) that EO data would benefit from rotation and translation invariant relative position encoding. This research framework—that uses intuition to combine existing methods in novel ways—is strongly represented in NeurIPS every year. Regarding comparisons, we extensively compare CROMA to all foundation models for EO. There are many recent frameworks invented in computer vision that we could leverage to pretrain models on the SSL4EO dataset—but doing so for all new approaches is not practical. Specifically, two concerns are raised: (i) we do not compare to a SatMAE model pretrained on SSL4EO, and (ii) we do not compare to a “SatMAE + contrastive learning” framework. Regarding the first concern, we do not believe that SatMAE pretrained on SSL4EO would improve SatMAE’s performance on benchmarks because, qualitatively, the data distribution of fMoW-Sentinel is closer to these benchmarks than SSL4EO (primarily, the sizes of images). In fact, CROMA outperforms SatMAE when finetuning on the fMoW-Sentinel dataset—this demonstrates that our framework learns better representations than SatMAE. Regarding the second concern, the “SatMAE + contrastive learning” framework does not exist, but we could try it. However, we do not expect it to outperform CROMA because CROMA outperforms VICRegL (see our appendix), which outperforms unimodal MAE + contrastive learning (see the VICRegL paper). We would have been happy to address these two concerns with experiments during the rebuttal period, but these concerns were not raised during the original review. Overall, we are disappointed that these new concerns dropped our score from a borderline accept to a borderline reject.

Reviewer iZ736/10 · confidence 5/52023-07-04

Summary

The paper presents a self supervised representation learning model for multimodal sentinel images. The model learns from geographically aligned optical and radar (sentinel-2 and sentinel-1, respectively) representations that are then used for downstream tasks, such as classification and segmentation. In the paper, classification and segmentation tasks are illustrated by using different approaches, namely, fine tuning, linear probing and nonlinear probing (MLP). Authors also show other quantitative evaluations (knn, kmeans over classes, a UMAP) to show the quality of the learned representation over SatMAE [26], which is the main competing method.

Strengths

- The paper deals with an important topic, and the models and results presented in the paper are significant for a variety of applications making use of sentinel data. - Results are validated on well known datasets in the field, and show superiority over a series of strong baselines and several metrics. Ablation study is very complete and shows how the model performs under different changes in the modules. - The approach of combining masked out reconstruction and contrastive losses in learning SSL representations is, to the best of my knowledge, novel in the field of geographically aligned data. The use of different modalities is also interesting, although the synergistic use of optical and radar images is well known and studied, but for specific applications. - The paper is very dense but clear enough, well written and well structured.

Weaknesses

I think that the paper is sound and although it touches upon a niche application of computer vision that might not be of wide interest for the NeurIPS audience, it could be a good contribution. However I have a series of comments that I think could be addressed and improve the paper. In general: - The data description and the explanation of the different levels of preprocessing for each dataset (eg atmospheric corrections from L1C to L2A), and how those influence the model, are not well explained. For instance, sentinel 2 has 13 channels, but of which only a subset are useful for land cover applications. Also, spatial resolution of the different S2 channels is very different (from 10 to 60m) and the one from S1 changes depending on the processing levels. - The different benchmark tested are characterised by different preprocessing and channel subsets, and it is unclear how the models are fine tuned in this setting and what is the dependency on those aspects. - I felt that the related work section at L72 makes plenty of references but it is not so good at clarifying some main lines of research, pros and cons of those. There are many papers, maybe too many, and it is hard to get some information out of that.

Questions

- L91: I found the concept of "optionally multimodal" very interesting, and I think it could be better framed. I am not sure whether the proposed CROMA it is indeed optionally multimodal, but the fact that some image pairs are not available concurrently is a very common issues which is often solved by temporal composition, but might not be optimal for some fine grained monitoring applications. - L109: it is not completely clear to me why (beyond the ablations) RPE offer, sometimes, good performance when extrapolating over larger images. At the same time, i think, contrarily what stated afterwards, the wild variability of remote sensing images is not an issue: images are referenced absolutely to geographical coordinates and can be cropped following the constraints of the task and hardware. - What is not so clear to me, is how the choice of spectral channels for S2 is done, and how resolution is dealt with. S2 channels have a ground sampling distance that varies between 10m, 20m and 60m, while radar can vary depending on the processing level. It is unclear how patches are sampled and how the mismatches in resolution is dealt with, since if sampling eg a 80x80 pixels patch, the actual content varies a lot unless interpolated and upsampled, which could introduce artefacts in itself. I think these data preprocessing steps should be better explained. - It is unclear at which level of processing the data is used, whether L2A corrected or L1C Sentinel-2 data, and if so, how the processing is performed. It is also mentioned that 12 bands are used, but in facts S2 has 13. Just that some of these bands are used for sensing properties of the atmosphere such as clouds and aerosols, and do not help in performing land-cover / land-use related modelling. Again, I think that some of these aspects should be clarified in the main text. - I wonder why [86] has not been included in the baselines, as it is one of the main papers highlighted in the related work section dedicated to RS representaiotns. - In light of the above points, some datasets used to test the model have very different properties and characteristics. how is the CROMA model pretrained on 12 S2 channels, retrained for each of the benchmark making use of the specific data provided? eg the fMoW, as far as i remember, only make use of 8 channel and not 12. - L203 and following: It is unclear why no other optical image participates in the definition of negative samples, this could be maybe beneficial to encode differences in landcover at different locations, or account for seasonality effects. - L240: it is unclear what "single label benchmarks" are.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

4 excellent

Contribution

2 fair

Limitations

- Limitations shortly mentioned in the conclusion section. I agree with those highlighted. - I had the feeling when reading the paper, the fact that CROMA was performing better than SatMAE, that is all it was needed. I think that SatMAE comes with pros and cons that are not very well discussed and framed. I think that since the paper is comparing and improving directly upon SatMAE, these aspects could have been better presented. - I really missed a test on non Sentinel data. I do agree that Sentinel is a great data source, but it is not the only source used, particular when studies need to go back before 2015/2016. It would have been nice to see some results on other satellite data, but I understand that this could have been too much work or out of scope, but I would nonetheless mention it (this goes beyond higher spatial and spectral resolution, it is often interesting to transfer models to lower spatial and spectral resolution).

Reviewer iZ732023-08-14

Follow up

I'd like to thank the Authors for the follow up. From my perspective, the rebuttal is clear, and definitely improves my understanding of the paper. There are still a couple of minor points that are not explicitly addressed, but I agree those are not worthy discussion at this stage and can be directly incorporated in the camera ready. I think the contribution is interesting, and overall relevant not only to the EO community, so I am happy to increase the score to weak accept, and looking forward to discuss further with other reviewers, if needed.

Reviewer vsHk2023-08-20

Thanks for the reply

Thanks the authors for their reply. Although parts of the issues have been modified, the novelty of this work is limited to NIPS. In addition to the poor organization and written, the contributions of this manuscript is narrowed. Thus, I insist on my original decision, i.e., Reject.

Reviewer FExC2023-08-20

Response

I would like to thank the authors for their time and the level of details provided to address my questions. In particular, thanks for the in-depth explanation of X-ALiBi's capability to be extended towards multiple modalities, not only coming from different (registered) sensors but also from different resolutions. This renders the presented approach to be more scalable wrt. to input modalities than published work in this area employing straightforward contrastive learning approaches for pre training. Given that my questions are answered and I think that this submission outlines a very interesting research direction to investigate, I decide to **increase** my rating.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC