Posterior Contraction Rates for Matérn Gaussian Processes on Riemannian Manifolds

Gaussian processes are used in many machine learning applications that rely on uncertainty quantification. Recently, computational tools for working with these models in geometric settings, such as when inputs lie on a Riemannian manifold, have been developed. This raises the question: can these intrinsic models be shown theoretically to lead to better performance, compared to simply embedding all relevant quantities into $\mathbb{R}^d$ and using the restriction of an ordinary Euclidean Gaussian process? To study this, we prove optimal contraction rates for intrinsic Matérn Gaussian processes defined on compact Riemannian manifolds. We also prove analogous rates for extrinsic processes using trace and extension theorems between manifold and ambient Sobolev spaces: somewhat surprisingly, the rates obtained turn out to coincide with those of the intrinsic processes, provided that their smoothness parameters are matched appropriately. We illustrate these rates empirically on a number of examples, which, mirroring prior work, show that intrinsic processes can achieve better performance in practice. Therefore, our work shows that finer-grained analyses are needed to distinguish between different levels of data-efficiency of geometric Gaussian processes, particularly in settings which involve small data set sizes and non-asymptotic behavior.

Paper

References (64)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer dGg78/10 · confidence 3/52023-06-28

Summary

The paper concerns Gaussian processes on manifolds. The authors present theorems for contraction rates for Matern Gaussian processes defined intrinsically and extrinsically through and embedding in a higher-dimensional Euclidean space. The authors shows that the rates are asymptotically equal in the two settings. Additionally, they treat the case of finitely truncated series expansion of the kernels to get a similar rate. Finally, it is shown experimentally that it can still be beneficial to work with the intrinsic geometry in the small sample domain.

Strengths

- extremely well-written paper. I found the exposition very clear - novel theoretical results - interesting findings (in line with what the authors state, I would not have expected the manifold dimension to show up in the extrinsic case) - empirical study to cover the small sample case

Weaknesses

- since the theorems are not in the main paper, one could perhaps consider if a longer format than NeurIPS (a journal paper) would be more suitable for the paper from a presentation point of view

Questions

no questions

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

yes

Reviewer woEb8/10 · confidence 5/52023-07-02

Summary

This paper studies the contraction rate(s) for both the intrinsic and extrinsic Mat\'ern Gaussian process in compact Riemannian manifold. The authors proved that the (optimal) rate in both cases is $\frac{2 \min(\beta, \nu)}{2 \nu +d}$, where $\nu$ is the smooth parameter of the Mat\'ern process, $\beta$ is the regression function smoothness class, and $d$ is the dimension. The authors also showed with examples that the geometric models outperform the non-geometric ones through empirical error analysis.

Strengths

The results in this paper are novel and enlightening. Up to minor typos, I enjoy reading the paper. It is concerned with the fine topic of Gaussian processes on manifolds, and on why (and how) this kind of modeling is valuable through quantitative posterior contraction analysis. The main contribution is the (optimal) contraction rate of both intrinsic and ambient Mat\'ern processes on compact manifolds in the nonparametric setting, the analysis of which is different from that in the Euclidean setting as the definition of the Mat\'ern processes on manifolds is subtle (though the rates are the same). It is also shown how the underlying geometric analysis outperforms by numerical experiments.

Weaknesses

As for every (good) paper, there are always plenty of things remained to be done. For instance, (1) In the current analysis it is assumed that the "nugget" $\sigma_\epsilon$ is given. In Bayesian setting, it would be interesting to know what happens if we put a prior on it (and more importantly, what prior to put). (2) The results, implicitly, compute the interpolation errors (as $L^2$ error based on $p_0$). What about extrapolation errors (that is, to predict an "outside" the domain)? This is also important in most geostatistics problems. (3) For the numerical experiments, the authors only consider the synthetic examples, e.g. dragon, sphere. It is interesting to see if the intrinsic and extrinsic modeling can bring a big difference for some real data. One such example is provided in [14].

Questions

While the paper is generally well written, I feel the authors need to address the following points: (1) I catch the idea of the intrinsic vs extrinsic modeling quickly. But the authors may point out a reference for this terminology (is this inspired from [15])? (2) p.3, line 80: I guess the authors mean $f \sim GP(m, k)$ instead of $GP(0,k)$. (3) p.3, line 115-116: the notation $C^{\beta}$ for the H\"older class is nonstandard. Does this mean $\mathcal{C}^{0, \beta}$? I guess not since $\beta > \frac{d}{2}$ meaning that it can be larger than $1$. The authors may need to clarify this. (4) In Result 1, it is assumed that $f$ is mean zero. Is this also assumed in Theorems 5 and 6? It is known that for the Mat\'ern processes in the Euclidean space, the mean function may raise the identifiability issue (see Stein's Interpolation of Spatial Data, or Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models). There are also related discussions in [24]. The authors may need to clarify, and add a few sentences in the paper. (5) p.4, Assumption 3: the authors may "indicate" $\sigma_\epsilon$ is known earlier, as this is important for the Bayesian workers. (6) Regarding the use of Theorems 5 and 6: it is clear from the rates that one should take $\nu = \beta$ (i.e. if one can identify the function class, and set the same smoothness parameter in the Mat\'ern process). However, we never know exactly $\beta$. The question is whether there is an adaptive way to select $\nu$. (Of course, this may be discussed in another paper. I only want to bring this question to the authors.) (7) Regarding Theorem 8: the rate is $\frac{2 \min(\beta , \nu)}{2 \nu +d} = \frac{2 \nu}{2 \nu +d}$ if $\beta = \nu$. This rate also appears in the prediction based on BLUE (best linear unbiased predictor) on the context of MLE. See e.g. Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models, JRSS-B, 2021, page 1055. The authors may want to mention this as well. (8) There are a few more references that the authors may want to add. (a) Stein's book Interpolation of Spatial Data is one of the main references on the Mat\'ern Gaussian processes. (b) Tang, Zhang and Banerjee's paper On identifiability and consistency of the nugget in Gaussian spatial process models studies the Mat\'ern process (with nugget) in the Euclidean space, where the identifiability issue occurs as pointed out in (7). This is related to the setting of Theorem 8. Arafat, Porcu, Bevilacqua and Mateu's paper Equivalence and orthogonality of Gaussian measures on spheres, Journal of Multivariate Analysis, 2018 is also relevant. (c) Regarding the prior on $\sigma^2_\epsilon$, one such model is the conjugate Bayesian linear model in Banerjee's paper Modeling massive spatial datasets using a conjugate bayesian linear modeling framework, Spatial Statistics, 2020; this model was analyzed in Zhang, Tang and Banerjee's Exact Bayesian geostatistics using predictive stacking, arXiv:2304.12414, Section 2. I believe it should be possible to generalize the results in this paper to the conjugate bayesian linear model.

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

NA

Reviewer qZfm7/10 · confidence 3/52023-07-31

Summary

This paper establishes bounds on the contraction rate of matern processes on Riemannian manifolds. The authors study three variants: 1) intrinsic matern process 2) truncated intrinsic matern process 3) extrinsic matern process and show that in each case, the optimal contraction rate can be achieved, which matches the Euclidean case.

Strengths

This is a fundamental problem, and it is remarkable that the authors are able to prove the same optimal contraction rates for both the intrinsic and extrinsic Riemannian matern processes. Furthermore, the manifold hypothesis has been receiving increasing attention recently, and so the analysis of Gaussian processes on manifolds is well motivated. I also appreciate the examples that the authors provided for illustrating the difference between intrinsic and extrinsic processes.

Weaknesses

see questions below

Questions

Regarding the smoothness parameter nu: 1) Can the authors provide intuition for why nu (as used in (7) and (10)) is the smoothness parameter (i.e. smoothness in what sense)? 2) The Riemannian matern process (7) and the Euclidean matern process (10) appear quite different, can the authors explain the analogy between these two processes, and specifically, why is nu, as used in (7), comparable to nu as used in (10)? In particular, can the authors explain the comment from line 275 "the intrinsic Matérn process, its truncated version and the extrinsic Matérn process all possess the same posterior contraction rates", and how these two rates are comparable under the different assumptions? 3) All theorems assume nu > d/2. Is this a standard assumption? How necessary is this assumption, and the theorems not work when nu is small? (and I have the a similar question for the beta > d/2 assumption in assumption 3 as well) Regarding the intrinsic matern process: 4) expression (7) involves an sum over eigenfunctions of the laplace beltrami operator. The authors do mention that this sum can be truncated, but for a general manifold, it seems like even computing a single eigenfunction can be quite expensive. Can the authors comment on how this is done computationally (and what is the cost), when the manifold is an arbitrary one, e.g. the dragon? 5) In figure 2, the authors give a dumbbell example which highlights the difference between intrinsic and extrinsic kernels -- one important difference seems to be that two points can be far away in manifold distance, but close in euclidean distance (under the embedding). Consequently, a function may have a small lipschitz constant wrt manifold distance, but a huge lipschitz constant wrt euclidean embedding distance. Intuitively, why is this not reflected in the contraction rates? Is it because the assumptions made do not care about things like lipschitz smoothness? (related to my earlier question of what is the meaning of nu?) Regarding the bound: 6) What is contained in the constant C in the theorems, is this a universal constant? Or does it depend polynomially/exponentially on any problem parameters?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

n/a

Reviewer yf8T8/10 · confidence 3/52023-08-01

Summary

This paper investigates the theoretical properties and performance of Gaussian processes in machine learning, particularly when applied in geometric settings such as Riemannian manifolds. It compares intrinsic and extrinsic methods, with the former directly formulated on the manifold of interest and the latter requiring a higher-dimensional Euclidean space embedding. The research derives posterior contraction rates for three primary geometric model classes and shows that all three can lead to optimal procedures, given certain conditions. Empirical experiments support these theoretical findings, demonstrating better performance by intrinsic models in small-data regimes.

Strengths

NA

Weaknesses

NA

Questions

NA

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

NA

Reviewer woEb2023-08-11

Thanks for the detailed explanations. The score remains unchanged.

Reviewer dGg72023-08-12

Thanks for the rebuttal. My scoring has not changed.

Reviewer qZfm2023-08-13

Thank you for the rebuttal. My questions are adequately addressed, and I have increased my score to 7.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC