Summary
This paper studies the contraction rate(s) for both the intrinsic and extrinsic Mat\'ern Gaussian process in compact Riemannian manifold. The authors proved that the (optimal) rate in both cases is $\frac{2 \min(\beta, \nu)}{2 \nu +d}$, where $\nu$ is the smooth parameter of the Mat\'ern process, $\beta$ is the regression function smoothness class, and $d$ is the dimension. The authors also showed with examples that the geometric models outperform the non-geometric ones through empirical error analysis.
Strengths
The results in this paper are novel and enlightening. Up to minor typos, I enjoy reading the paper. It is concerned with the fine topic of Gaussian processes on manifolds, and on why (and how) this kind of modeling is valuable through quantitative posterior contraction analysis. The main contribution is the (optimal) contraction rate of both intrinsic and ambient Mat\'ern processes on compact manifolds in the nonparametric setting, the analysis of which is different from that in the Euclidean setting as the definition of the Mat\'ern processes on manifolds is subtle (though the rates are the same). It is also shown how the underlying geometric analysis outperforms by numerical experiments.
Weaknesses
As for every (good) paper, there are always plenty of things remained to be done. For instance,
(1) In the current analysis it is assumed that the "nugget" $\sigma_\epsilon$ is given. In Bayesian setting, it would be interesting to know what happens if we put a prior on it (and more importantly, what prior to put).
(2) The results, implicitly, compute the interpolation errors (as $L^2$ error based on $p_0$). What about extrapolation errors (that is, to predict an "outside" the domain)? This is also important in most geostatistics problems.
(3) For the numerical experiments, the authors only consider the synthetic examples, e.g. dragon, sphere. It is interesting to see if the intrinsic and extrinsic modeling can bring a big difference for some real data. One such example is provided in [14].
Questions
While the paper is generally well written, I feel the authors need to address the following points:
(1) I catch the idea of the intrinsic vs extrinsic modeling quickly. But the authors may point out a reference for this terminology (is this inspired from [15])?
(2) p.3, line 80: I guess the authors mean $f \sim GP(m, k)$ instead of $GP(0,k)$.
(3) p.3, line 115-116: the notation $C^{\beta}$ for the H\"older class is nonstandard. Does this mean $\mathcal{C}^{0, \beta}$? I guess not since $\beta > \frac{d}{2}$ meaning that it can be larger than $1$. The authors may need to clarify this.
(4) In Result 1, it is assumed that $f$ is mean zero. Is this also assumed in Theorems 5 and 6? It is known that for the Mat\'ern processes in the Euclidean space, the mean function may raise the identifiability issue (see Stein's Interpolation of Spatial Data, or Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models). There are also related discussions in [24]. The authors may need to clarify, and add a few sentences in the paper.
(5) p.4, Assumption 3: the authors may "indicate" $\sigma_\epsilon$ is known earlier, as this is important for the Bayesian workers.
(6) Regarding the use of Theorems 5 and 6: it is clear from the rates that one should take $\nu = \beta$ (i.e. if one can identify the function class, and set the same smoothness parameter in the Mat\'ern process). However, we never know exactly $\beta$. The question is whether there is an adaptive way to select $\nu$. (Of course, this may be discussed in another paper. I only want to bring this question to the authors.)
(7) Regarding Theorem 8: the rate is $\frac{2 \min(\beta , \nu)}{2 \nu +d} = \frac{2 \nu}{2 \nu +d}$ if $\beta = \nu$. This rate also appears in the prediction based on BLUE (best linear unbiased predictor) on the context of MLE. See e.g. Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models, JRSS-B, 2021, page 1055. The authors may want to mention this as well.
(8) There are a few more references that the authors may want to add.
(a) Stein's book Interpolation of Spatial Data is one of the main references on the Mat\'ern Gaussian processes.
(b) Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models studies the Mat\'ern process (with nugget) in the Euclidean space, where the identifiability issue occurs as pointed out in (7). This is related to the setting of Theorem 8. Arafat, Porcu, Bevilacqua and Mateu's paper Equivalence and orthogonality of Gaussian measures on spheres, Journal of Multivariate Analysis, 2018 is also relevant.
(c) Regarding the prior on $\sigma^2_\epsilon$, one such model is the conjugate Bayesian linear model in Banerjee's paper Modeling massive spatial datasets using a conjugate bayesian linear modeling framework, Spatial Statistics, 2020; this model was analyzed in Zhang, Tang and Banerjee's Exact Bayesian geostatistics using predictive stacking, arXiv:2304.12414, Section 2. I believe it should be possible to generalize the results in this paper to the conjugate bayesian linear model.
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.