Validating Climate Models with Spherical Convolutional Wasserstein Distance

The validation of global climate models is crucial to ensure the accuracy and efficacy of model output. We introduce the spherical convolutional Wasserstein distance to more comprehensively measure differences between climate models and reanalysis data. This new similarity measure accounts for spatial variability using convolutional projections and quantifies local differences in the distribution of climate variables. We apply this method to evaluate the historical model outputs of the Coupled Model Intercomparison Project (CMIP) members by comparing them to observational and reanalysis data products. Additionally, we investigate the progression from CMIP phase 5 to phase 6 and find modest improvements in the phase 6 models regarding their ability to produce realistic climatologies.

Paper

Similar papers

Peer review

Reviewer vRQU5/10 · confidence 4/52024-07-11

Summary

The paper proposes a new distance measure based on Wasserstein distance for data on a sphere. The work applies the methodology to climate model data, with primary focus on ranking climate models based on their agreement with reanalysis data.

Strengths

The paper is well-written and easy to follow. It provides adequate background discussion on both, the methodology and the specific problem of climate model inter-comparison. The methodology is introduced rigorously, carefully defining the terms and the associated spaces. The experiment section includes a large number of models (which is not a small undertaking given the size of climate simulation data).

Weaknesses

While the specific methodology is novel and may be of interest beyond climate modelling, it is a minor extension to the existing methods. Furthermore, the paper is heavily focused on a specific application. Consequently, a venue that is primarily focused on climate informatics would be more appropriate.

Questions

From what I understand, the data used in daily averages of temperature and precipitation for a historic period. Is the distance calculated on the daily basis? If so, how do you take into account the fact that climate models are known to be poorly temporally aligned with observation data? Would it make more sense to aggregate the data temporally? If so, do you have any thoughts on optimal ways to do that (i.e. how to pick the optimal window for aggregation)? You mention that baseline methods (e.g. RMSE) are unable to detect the variance of the anomalies. Do you have any thoughts on how your proposed methodology performs in the tails of the distributions that you're comparing? In most situations, the only parts of the climate distributions that are of interest to the end users are the extremes. Do you have any thoughts on whether it makes sense to rank climate models in the first place? I assume the ranking can then be used to create a weighted ensemble of models. Minor: Line 52: for such a purpose Fig. 2 y-axis - preCipitation

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

1

Limitations

Both, methodological and application specific limitations are discussed in Sec. 5.

Reviewer qRTk7/10 · confidence 4/52024-07-12

Summary

The paper defines SCWD as a special case of their proposal for functional sliced WD which it uses to compare CMIP members against reanalysis data. Additionally, with this new distance, it analyses the effectiveness of CMIP phase 6 over phase 5

Strengths

- The paper presents its ideas succinctly - Motivates the need to find a good distance measure for comparing distributions of functions defined over $L^2(S^2)$ - Provides a smooth transition from sliced WD to its functional variant - Method seems robust to kernel parameter - Allows easy visualization of the differences and helps isolate regions where the fields diverge - Validation experiments were sufficiently extensive

Weaknesses

- Would have been interesting to see VAEs as a baseline as proposed by [1] [1] Mooers, G., Pritchard, M., Beucler, T., Srivastava, P., Mangipudi, H., Peng, L., ... & Mandt, S. (2023). Comparing storm resolving models and climates via unsupervised machine learning. Scientific Reports, 13(1), 22365.

Questions

- L258 suggests that the resolutions between multiple outputs isn't consistent throughout? Have they been readjusted to similar spatial dimensions?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes

Reviewer UJm57/10 · confidence 5/52024-07-12

Summary

The paper introduces a new method for validating climate models by comparing their outputs to reanalysis data. The proposed method, Spherical Convolutional Wasserstein Distance (SCWD), accounts for spatial variability and local differences in the distribution of climate variables. The authors apply SCWD to evaluate historical model outputs from the Coupled Model Intercomparison Project (CMIP) phases 5 and 6, demonstrating modest improvements in the phase 6 models in producing realistic climatologies. The technical claims are well-supported by thorough theoretical and empirical analyses. The authors provide a robust mathematical foundation for SCWD and demonstrate its effectiveness in capturing spatial variability through extensive experiments. The paper is well-structured and written, with detailed explanations of the methodology and comprehensive evaluation results. However, some sections could benefit from additional clarity, particularly the mathematical derivations and kernel selection process.

Strengths

Originality: The introduction of SCWD as a new metric for climate model validation is innovative and addresses the limitations of existing methods. Quality: The methodology is rigorously developed and supported by extensive experimental validation using real-world climate data. Clarity: The paper provides clear explanations of the SCWD methodology, supported by visualizations and detailed examples. Significance: The proposed method has significant implications for improving climate model validation, which is crucial for accurate climate projections and policy-making.

Weaknesses

Mathematical Derivations: Some mathematical derivations, particularly those related to the convolution slicer and kernel functions, could be explained more clearly to enhance understanding. Generalization: While the method is well-validated on historical climate data, additional experiments on different climate variables and temporal resolutions would strengthen the generalizability of the findings.

Questions

Could the authors provide more details on the selection process and theoretical justification for the specific kernel function used in SCWD? Have the authors considered applying SCWD to other climate variables or different temporal resolutions to evaluate its generalizability?

Rating

7

Confidence

5

Soundness

4

Presentation

3

Contribution

4

Limitations

The authors adequately address the limitations of their work, including the need for device-aware optimizations and the challenges in parameter selection. They also acknowledge the potential for further improvements and generalization of SCWD, providing constructive suggestions for future research.

Reviewer V9JK8/10 · confidence 3/52024-07-16

Summary

Developing metrics for comparison between high dimensional, multivariate climate models is an important and open area of study . Vissio et al (2020) proposed the use of the Wasserstein distance to quantify the similarity between climate models. However this approach involves spatial averaging, and therefore significant information is lost. This work introduces functional sliced Wasserstein distance in spherical coordinates, which provides a computationally tractable Wasserstein metric without spatial averaging. The method is demonstrated by comparisons between CMIP model data and ERA reanalysis data.

Strengths

The work is timely and presents a strong contribution to the field of climate science. The presentation is excellent - motivation and connections to previous work are clearly established. The new method is clearly explained, and demonstrated in a sensible set of experiments. The capability of the slicing kernel to focus on specific local regions provides a tremendous amount of flexibility to the metric, which will have utility in a wide range of important applications. Comparisons to other standard metrics are also made, and in cases where there are discrepancies with baselines, these discrepancies are discussed. Finally the authors speculate on potential applications beyond climate science.

Weaknesses

The paper has no obvious weaknesses.

Questions

I am wondering whether the method could be demonstrated on a simpler toy problem where the ground truth is better established, and the complexity and high dimensionality of the system is retained, before application to a reanalysis-vs-model comparison. As discussed in section 4.1, reanalsyis comparisons are still subject to discrepencies from other factors such as model physics.

Rating

8

Confidence

3

Soundness

4

Presentation

4

Contribution

3

Limitations

Limitations of both the method (lines 107-119 and 356-359) and the results (section 4.1) are discussed.

Authorsrebuttal2024-08-07

Visibility of global rebuttal

Hello, we want to make sure that the reviewers are able to see our global rebuttal. On our end, it seems that they do not have reading permissions, and we were unable to add these permissions when submitting the rebuttal. We can repost that content using official comments if necessary! Thank you.

Reviewer vRQU2024-08-10

Thank you for the thorough response. I particularly appreciate the comments on suitability for the venue, and I will increase my original score. However, I still feel that the paper would have a greater impact in a more climate-focused venue.

Reviewer qRTk2024-08-11

Thank you for the response and for clarifying how you handled regridding. I find this paper particularly valuable to climate science community, and I will increase my original score.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC