Pointwise uncertainty quantification for sparse variational Gaussian process regression with a Brownian motion prior

We study pointwise estimation and uncertainty quantification for a sparse variational Gaussian process method with eigenvector inducing variables. For a rescaled Brownian motion prior, we derive theoretical guarantees and limitations for the frequentist size and coverage of pointwise credible sets. For sufficiently many inducing variables, we precisely characterize the asymptotic frequentist coverage, deducing when credible sets from this variational method are conservative and when overconfident/misleading. We numerically illustrate the applicability of our results and discuss connections with other common Gaussian process priors.

Paper

References (42)

Scroll for more · 30 remaining

Similar papers

Peer review

Reviewer oD8s6/10 · confidence 3/52023-06-15

Summary

Broadly, the paper addresses the issue of the quality of the approximate posterior of a sparse GP model in terms of uncertainty qualification (UQ). The paper chooses the setup with eigenvector-inducing variables and rescaled Brownian motion prior. However, as shown, the results also extend to other kernels (squared exponential (RBF) and Matérn kernels). Precisely, the paper's main conclusion is that with a well-calibrated prior and a sufficiently large number of inducing variables, the uncertainty quantification (UQ) obtained is reliable, though it can be conservative. The paper presents theoretical results on the rate and quality of SVGP posterior, nature of credible sets, and contraction rate of approximate posterior. The theoretical results are shown to match the empirical results based on experiments on simulation data and semi-simulation data.

Strengths

The topic of the paper is relevant as Gaussian processes are go-to models for many tasks, including Bayesian optimization, active learning, and various sequential tasks. Uncertainty quantification is one of the crucial properties of the Gaussian process models that make them popular for these tasks. Recently, a lot of research has been done on the quality of the approximate posterior of sparse GP models, and this paper rightly fits in there in terms of evaluating the quality of approximate posterior in terms of uncertainty quantification. The claims of the paper are sensible, and the detailed derivations in the appendix also look logical (I would like to see what other reviewers think about the proofs as that is not my core area). Apart from a few places, the paper's notations and flow are good and convenient to follow. The experiments are also chosen wisely and demonstrate the theoretical results.

Weaknesses

A couple of weaknesses that I find in this paper are: * The paper talks about sparse variational Gaussian process models (SVGP) throughout, but actually, the model considered in the paper is sparse Gaussian process regression (SGPR). It can be clearly seen from Eq (3). It is better to make this clear in the paper. With a non-Gaussian likelihood, the equations for the posterior are different from those presented here. * Section 2.1 has issues with the notations. - As pointed out earlier, the equations are of SGPR and not SVGP. - L104 mean of the GP `m` is a typo? - L117-L118 starts writing cov. Should it be `k`? - $K$ is the gram matrix and should be bold as I believe the paper tries to follow that notation by making $k_n(x)$ bold. * A key takeaway message of the paper is missing. There are a bunch of theorems, derivations, and contributions, but I am still trying to figure out the main takeaway message from the paper that I should remember next time I plan to use a sparse GP model. * As we discuss uncertainty quantification, does it make sense to report negative log predictive density (NLPD) in the experiments along with the RMSE? * A plot of the true posterior, approximate posterior, and credible set would be good to demonstrate the setup and the output. * Theorem 3.5 derives a bound for the number of inducing variables `m`. How does it relate to Burt et al. (2019), as they also derive a bound for the number of inducing variables?

Questions

* Is the metric for the quality of the approximate posterior different from the metric of UQ? Is having a “good” quality approximate posterior but not a well-calibrated UQ possible? As the authors also point out in the paper, there has been a lot of recent research evaluating the quality of approximate posterior and convergence rates. Is it much different? * What are the connections with Burt et al. (2019), and where do the authors agree and disagree with them? In general, I believe there should be more connections drawn with them. * A minor comment, references to the appendix need to be included. I do not see any reference to the appendix in the main paper, so connecting derivations and their corresponding sections becomes tricky. * More practical question: how to select the value of $\alpha, \gamma$ for a real-world dataset, let us say a UCI data set. I believe more evaluations should be done by extending section 4. It would make the work more applicable, in my opinion. * More questions are written in the weakness section.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

The authors do not discuss them, and I do not see a limitation directly of this work.

Reviewer sbVc7/10 · confidence 3/52023-07-03

Summary

The paper gives some theoretical results for frequentist coverage of the credible sets for sparce variational Gaussian processes using a class of Brownian motion based priors. It provides some conclusions relating the smoothness of the prior vs the smoothness of the true underlying function and the resulting ability of the Bayesian inference to produse reliable frequentist uncertainty estimates.

Strengths

The paper is very well written and easy to follow at a high level. The theoretical results are summarised well and give a good indication of the overall contribution of the paper. The contribution seems to be an important starting point for providing guarantees for pointwise uncertainty estimates in SVGPs. It gives enough details to indicate the exact extent to which the results are valid. The theoretical results are provided with a sufficient level of detail to be accessible by a more general (less statistically inclined) audience. The discussion provides a number of directions for future work that can be taken up by the community.

Weaknesses

The main motivation for the work is the wide use of SVGPs in practice. However, the results in the paper consider only very rough (BM) priors which are rarely used in practice, and only touch on smoother, more commonly used options such as Matern kernels. Admittedly, it is still an important contribution that further work can build on.

Questions

It would be useful for my own understanding if the authors commented more on the statement that $L_2$-type credible sets are less reflective of actual practice than pointwise credible sets. I would also be curious why the authors chose this venue for publishing this work, given the theoretical nature of the results.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

It is possible that the paper will not be very accessible to the broader audience that it is addressing due to the theoretical nature of the main results. That said, I think the authors did a good job highlighting the exact contribution in the introduction and it should be accessible for practitioners in informing them of the possible consequences in the estimates of credible intervals based on the choice of priors. As the choice of prior is typically difficult in practice, these types of results are valuable in practice.

Reviewer BMyX6/10 · confidence 3/52023-07-03

Summary

The paper gives algorithms for point wise estimation and uncertainty quantification (error bars) for sparse variational Gaussian processes with the Brownian motion as prior. The SGVP model though widely used in practice, lacks theoretical guarantees and this paper is the first to establish error guarantees for the model with Brownian priors. Experiments are presented to validate the results. The priors for SVGP models used in practice are not Brownian, extending the results in this paper to these priors remains an open question.

Strengths

- This paper is among the first to address rigorously the statistical inference properties of the low rank approximations used in Gaussian processes. The main result establishes rigorous bounds on the size of the confidence interval for the SGVP prediction given bounds on the smoothness of the prior and the ground truth and the number of eigenvalues retained in the approximation of the kernel matrix. This is a new and useful result for SGVPs. - The analysis is particular to the case of Brownian motion priors. BM priors have eigenvectors that can be expressed in closed form (they correspond to the trigonometric functions) and have strong self-similarity properties. However, in the experiments it is observed that the qualitative behavior that is observed for the BM priors is also observed for the Matern and squared exponential kernels used in practice.

Weaknesses

- The result does not extend in a straightforward way to the Matern and squared exponential kernels used in practice with SGVPs. Brownian priors have many special properties that are not shared with other Gaussian processes used in ML. - The result although a useful addition to the literature on SGVPs do not seem not be of enough general interest for a clear accept and need further development for analyzing the instances of SGVPs used in practice. - The results are non adaptive as they require a pre-specified prior smoothness which is not a realistic assumption in practice.

Questions

Pg 2, the wedge notation for the minimum, the definitions of smoothness and rate should be introduced before stating main results. It may be helpful to look at the stochastic process literature where the various properties of BM were originally developed, there may be other stochastic processes like the Ornstein-Uhlenbeck process considered in the stochastic process literature whose spectra are similar to the BM, however these do not seem to have not been used for SGVP.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Limitations are discussed clearly by the authors.

Reviewer tkgm4/10 · confidence 4/52023-07-05

Summary

This paper considers the problem of uncertainty quantification for SVGP models with a Brownian motion prior for a non-parametric estimation problem. Theoretical results are presented that show

Strengths

The results are fundamental in nature and of broad interest, and clearly explained except in places where quantities are not clearly defined.

Weaknesses

The main contributions of the paper are a little unclear. It is a bit unclear which results are known and which results are new. It is also unclear what is special about the rescaled Brownian motion prior. Does the paper recommend using the rBM prior over other priors because of some special properties. Otherwise the choice of rBM in the title is not well motivated. Does the proof technique also extend to more general prior models beyond rBM. The choice of rBM prior as the main focus is also not well motivated. If this is a standard choice, this must be explained clearly. It would perhaps be useful to provide a brief sketch of the proof. For eg the discussion below Proposition 3.2 is somewhat of a proof sketch but isn't labeled as such.

Questions

What precisely is Q*. Can you please define it formally in section 2. Perhaps Q* should include a gamma in the subscript. Is Q* the eigenvector SVGP or any SVGP as defined above (3). Similarly please also provide the definition of rescaled BM in the preliminaries. A formula for the Mattern kernel in Section 5 would be useful. Table 1 columns currently labeled m* and n need to be renamed to something more informative and which doesnt require reading the paragraph above the table before reading the table. The caption can also be more informative (and better spaced from the table). The explicit value of m* would also be useful to know instead of the formula in the caption. A brief description of coverage, length and RMSE is necessary to read the table clearly. Alas, these are relegated to the appendix. (Similarly for Table2) n,2 subscript in Proposition 3.4 is unclear. Is Lemma 3.1 a noteworthy result or can it be part of section 2? Similarly the paragraph above Prop 3.2 In Section 2.1, is it GP(nu0, k) instead of m? please dont use m for mean function since m is the number of inducing points. The space in which u_i belong is quite unclear. It appears they are actually part of the target space (ie same as space of y_i). Then the choice of the naming them Eigenvector inducing features is perhaps misleading, and Eigenvector inducing "targets" is perhaps more appropriate. The term cov(f(x), u_i) is confusing. How is this covariance function defined? since the kernel requires two points in the domain as input. Please clarify this immediately following the definitions of Kxm etc., rather than point the reader to [38]. The bolding can be more consistent. I think r_m is a vector which is not bolded unlike other vector quantities. In Section 3.1 end of line 150, shouldn't this also be f(x)|Dn since SVGP also is conditional on n samples? The choice of naming tn as "frequentist variance" is unclear. Maybe a comment on that (at least in OpenReview) would be useful, but also to the reader of this paper. As suggested earlier, the discussion leading upto Prop 3.2 can be part of Section 2 since everything follows naturally from the preliminaries and does not require any special mention in the section on Main results. The role of Mn in all results is a little unclear. An example Mn and its implications would be useful for the reader.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

No limitations envisioned.

Reviewer oD8s2023-08-11

Thank you for the response! > As we discuss uncertainty quantification, does it make sense to report negative log predictive density (NLPD) in the experiments along with the RMSE? We have added this to the simulations, thank you. Is it possible to share the numbers/table with RMSE and NLPD here?

Authorsrebuttal2023-08-12

NLPD Values

Hello! Unfortunately I wasn’t able to get the latex tables to compile here, but here are the values for the NLPD corresponding to the cases presented in the original version (of course in our updated document we include also the coverage, length and RMSE, but doing so here makes it quite difficult to read. Note here that the column SGPR represents the NLPD for the variational posterior, and GP the NLPD for the normal posterior. ### Table 1 Prior | alpha | gamma | SGPR | GP —————————— —————— **Fixed Design** —————————— —————— rBM | 1.0 | 0.5 | -0.90 | -0.90 Mat | 1.0 | 0.5 | -0.68 | -0.68 SE | 1.0 | 0.5 | -0.67 | -0.67 —————————— —————— rBM | 0.5 | 0.5 | -0.09 | -0.09 Mat | 0.5 | 0.5 | -0.40 | -0.40 SE | 0.5 | 0.5 | -0.21 | -0.21 —————————— —————— **Random Design** —————————— —————— rBM | 1.0 | 0.5 | -0.65 | -0.65 Mat | 1.0 | 0.5 | -0.55 | -0.55 SE | 1.0 | 0.5 | -0.02 | -0.02 —————————— —————— rBM | 0.3 | 0.5 | 2.23 | 2.23 Mat | 0.3 | 0.5 | 0.91 | 0.91 SE | 0.3 | 0.5 | 0.36 | 0.36 —————————— —————— ### Table 2 **Multidimensional Random Design** Type | rho |SGPR | GP —————————— —————— Uniform | n/a | 0.96 | 0.96 Gaussian | 0.0 | 1.34 | 1.34 Gaussian | 0.2 | 1.90 | 1.90 Gaussian | 0.5 | 1.36 | 1.36 —————————— —————— **Semi-synthetic Data** —————————— —————— SGPR | GP 3.22 | 3.22 Thanks.

Reviewer oD8s2023-08-14

Thanks for sharing the table. I am raising my score! I believe a discussion on extending the method to more priors (RBF, Matern, etc.) and an elaborate section on connection with Burt *et al.* (2019) would strengthen the work. Also, while reporting the numbers NLPD or RMSE, kindly report the standard deviation as well by performing K-fold cross-validation.

Reviewer sbVc2023-08-18

Thank you for the response and clarification.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC