Variational Gaussian Processes For Linear Inverse Problems

By now Bayesian methods are routinely used in practice for solving inverse problems. In inverse problems the parameter or signal of interest is observed only indirectly, as an image of a given map, and the observations are typically further corrupted with noise. Bayes offers a natural way to regularize these problems via the prior distribution and provides a probabilistic solution, quantifying the remaining uncertainty in the problem. However, the computational costs of standard, sampling based Bayesian approaches can be overly large in such complex models. Therefore, in practice variational Bayes is becoming increasingly popular. Nevertheless, the theoretical understanding of these methods is still relatively limited, especially in context of inverse problems. In our analysis we investigate variational Bayesian methods for Gaussian process priors to solve linear inverse problems. We consider both mildly and severely ill-posed inverse problems and work with the popular inducing variables variational Bayes approach proposed by Titsias in 2009. We derive posterior contraction rates for the variational posterior in general settings and show that the minimax estimation rate can be attained by correctly tunned procedures. As specific examples we consider a collection of inverse problems including the heat equation, Volterra operator and Radon transform and inducing variable methods based on population and empirical spectral features.

Paper

References (78)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Sbdt5/10 · confidence 3/52023-06-30

Summary

This paper analyzes posterior contraction rates for the inducing variables variational Bayes approximation for Gaussian process. In particular, mildly and ill-posed inverse problem settings, which are characterized by profiles of eigenvalues of measurement operators, are considered. For the both cases, the posterior contract rates are assessed theoretically. The usefulness of the results are shown for applications to three example problems.

Strengths

The posterior contract rates are assessed analytically for the inducing variable variational Bayes approximation of Gaussian process. This is useful to determine the optimal number of inducing variables in the approximation.

Weaknesses

The analysis relies on the assumption that the functional space where the target function belongs is known. However, there are many practical situations where such an assumption does not hold. The analysis is not applicable for such situations.

Questions

What happens if the target function does not belong to the assumed functional space?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Yes.

Reviewer KBMK6/10 · confidence 2/52023-07-03

Summary

This paper presents a theoretical analysis of the posterior contraction rates for the variational posterior in the context of linear inverse problems. Specifically, the variational posterior is obtained using the sparse Gaussian processes approach, with inducing variables learned through KL divergence minimization. The authors investigate three types of operators: the Volterra operator, the heat equation, and the Radon transform, each exhibiting varying degrees of ill-posedness. The numerical analysis focuses on synthetic toy problems, and the authors demonstrate that in this case, even a small number of inducing points can yield highly accurate results, achieving significant speed-up.

Strengths

[Originality] I am not up-to-date with theoretical results with frequentist-Bayes analysis, and how far the presented results are based on or deviates from the monograph Nickl, Richard, 2023. Nickl, Richard. "Bayesian non-linear statistical inverse problems." (2023). [Quality] The theoretical results seems sound. [Clarity] The paper is well-written; however, the use of jargon limits its accessibility to a narrower audience. [Significance] The inducing-point of Gaussian processes is widely used, so the theoretical results concerned with it should be of interest to the community working on inverse problems.

Weaknesses

[Quality] The experimental evaluation is limited to a synthetic toy example only. [Significance] The authors motivate the variational Bayes approach for computational reasons. However, recent developments of scalable GP methods can reduce the complexity of exact GP from O(n^3) to O(n^2) and effectively leverage the modern computing hardware such as GPU (Gardner, Jacob, et al. 2018). In this regard, one can afford many more inducing points in practice. Gardner, Jacob, et al. "Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration." *Advances in neural information processing systems* 31 (2018). The paper is a theory paper, but its practical significance appears to be lacking. It would greatly enhance the paper's value if the authors could at least mention some real-world applications in which the studied inverse problem commonly arises. For instance, discussing the application of the Radon transform in computed tomography (CT) could serve as a relevant and illustrative example.

Questions

Figure 1 shows that using m=3 is insufficient, while m=6 yields highly accurate results for the studied case, while this is a single case not quite informative. I am wondering whether it is possible to establish a "phase-transition curve" that can inform practitioners about the required number of inducing points for various types of inverse problems and different numbers of measurements. If the results are not understood correctly, it might leads to poor choice on the number of inducing points, and over-or-under estimate of the posterior uncertainty. This could be harmful in mission-critical applications such as medical imaging. How to prevent this happen? Could you please further explain why the posterior uncertainty being over-estimated while in traditional variational approximation the uncertainty usually got underestimated? How well is the uncertainty calibrated?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The authors do have discussed to what extend the presented results can be extended to nonlinear cases and other types of inverse problems.

Reviewer VbCA5/10 · confidence 2/52023-07-06

Summary

This paper analyses posterior contraction rates in Variational GP for mildly and severely ill-posed linear inverse problems. The authors prove that under certain conditions variational posterior achieves same contraction rates as the true posterior. Moreover, they show that given enough inducing points (in spectral domain) the above-mentioned condition holds. As specific examples, the minimal number of inducing points needed for optimal contraction and contraction rate is derived for two mildly ill-posed problems (Volterra operator, Radon transformations) and one severely ill-posed problem (heat equation).

Strengths

Analysis of posterior contraction rates improves our theoretical understanding of VGP for linear inverse problems. The presentation is well structured. The contribution is clear and well argued for.

Weaknesses

The experimental part in the main manuscript is very limited, with most experiments taken out to the appendix. I believe the paper is more suited for a journal format. It is very technical and requires a thorough review which is hardly possible in a conference format.

Questions

Suggestions: 1. Increase readability of figure 1. The legend way too small.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

The limitations are adequately addressed.

Reviewer TgVX7/10 · confidence 3/52023-07-25

Summary

This theoretical paper addresses the problem of inferring a function $f$ when $y(x_i) = \[Af\](x_i) + \epsilon_i$ with $i=1, ...,n $ and $A$ is suitably defined linear operator with $\epsilon_i$ being i.i.d Gaussian noise. In words this is a linear inverse problem. A (standard) Bayesian approach is adopted by putting a Gaussian process prior on $f$ with a carefully chosen covariance function. Then an interdomain variational inducing point approach is adopted. All this so far is standard and pre-existing. The contribution of the paper is to derive posterior contraction rates for this setup under certain assumptions and to apply it some specific examples with different choices of the the linear operator $A$. Specifically these are a volterra operator, a diffusion equation and a Radon transform.

Strengths

This is an important and challenging topic. The standard regression case has been studied by Nieman et al. 2022 which corresponds to taking the $A$ to be the identity in this paper. The extension is certainly of interest since the Bayesian approach to inverse problems is a thriving and important area. The required extension is treated well by the authors. The particular challenges of introducing an operator $A$ are clearly described and handled. The results derived require significant technical heavy lifting. I was not able to go through the proof in detail but the results and approach look sensible and I suspect it to be correct. The examples equations chosen were interesting and informative. The numerical simulations, if brief, were nonetheless adequate in a paper of this type and compelling. Limitations were communicated where they arose.

Weaknesses

**The paper requires strong assumptions on the kernel of the Gaussian process.** To some extent these are forced by theory since a very mispecified covariance function would not contract. But, as the authors clearly concede, there is a more general set of covariance functions that would work than the ones considered. **The two methods of choosing the inducing variables are somewhat idealized.** For a mildly ill-posed problem and the empirical spectral feature method the complexity is $O(N^2 M)$. Since $M$ the number of inducing points typically needs to be an increasing function of $N$ the number of data points then this is not much of an improvement on the exact inference which is $O(N^3)$ in practice and $O(N^{2.3})$ in theory. The authors do mention this. Alternatively the population method of choosing inducing variables uses $O(NM^2)$ but requires oracle knowledge of the eigensystem of $A^*A$. Admittedly the authors do show some interesting examples where they are able to do this. The required $M$ is more benign for the strongly ill-posed system and correspondingly the heat equation example given is quite compelling. **The inducing inputs are not optimized** This is known to make a big difference in the more practical part of the literature particularly in higher input dimensions. It is also one of the things that distinguishes the Titisias variational approach from the Deterministic Training Conditional (DTC) approach using the nomenclature of this paper https://www.jmlr.org/papers/volume6/quinonero-candela05a/quinonero-candela05a.pdf. **The literature review is not complete.** There are two significant papers from Pinski et al. Specifically: *Kullback--Leibler Approximation for Probability Measures on Infinite Dimensional Spaces. Pinski, Simpson, Stuart and Weber. 2015* *Algorithms for Kullback--Leibler Approximation of Probability Measures in Infinite Dimensions. Pinski, Simpson, Stuart and Weber. 2015* These papers are early, rigorous and are in a similar area to the submission. The method is distinct since they work principally with the inverse covariance operator. Importantly they also address the challenging non-linear case. The type of guarantee derived is also different since they do not derive contraction directly. Instead in the second they use their approximating for an MCMC proposal and gain asymptotic guarantees that way. Nonetheless these papers should be discussed. A historical note. The work of *Burt et al. 2019* is excellent and highly impactful but perhaps a bit too much is attributed to it here. It provides *rates* of convergence to the posterior but there was some work on convergence before this. This is largely because convergence follows straightforwardly from the variational formulation. *Titsias 2009* (the longer version) Section 2.3 has some discussion of improvement in terms of KL. This was formulated in terms similar to those used in the present submission by *Matthews et al. AISTATS 2015 On Sparse variational methods and the Kullback-Leibler divergence between stochastic processes.* The AISTATS paper mentioned also references several earlier foundational papers and the present submission would benefit from mentioning them too. For instance the early work authored by *Opper*, *Seeger* and *Csato* amongst others. None of these existing papers removes the main contribution of the current work. **Smaller points/ typos:** For the mildly ill-posed case described in Definition 1 shouldn't the decay be polymonial rather than exponential? On line 346 I believe the units need to be milliseconds for the description to make sense. So I think the typo is that "50.5 m" should be "50.5 ms".

Questions

I have no questions other than the ones mentioned above.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

There are no negative impacts emerging from this work that I can see. Limitations are adequately discussed in the submission for the most part. The "weaknesses" section of this review has further discussion.

Reviewer rih96/10 · confidence 3/52023-07-26

Summary

The authors propose an inducing-point variational Bayes approximation for linear inverse problems. They study the frequentist properties of the arising variational posterior approximation. Specifically, they prove contraction rates for the variational posterior, both in the mildly and severely ill posed case. Moreover, they illustrate the practical relevance of their results with a number of synthetic data experiments.

Strengths

The paper is well-written and its objectives and scope are clear. It provides an important contribution. As far as I am aware, frequentist properties of variational posteriors arising as inducing-point approximations in linear inverse problems have not been studied explicitly before. Such properties and contraction rates are useful in that they illustrate how the choice of the prior and the number of inducing points in the approximation impacts the statistical properties of the variational posterior. In that sense, they can provide useful guidance for the user.

Weaknesses

I am not exactly sure if the comparison to the existing literature provided by the authors is substantial enough – I address this among my questions to the authors.

Questions

In 2020, three papers were published in the Annals of Statistics providing a comprehensive and general study of the frequentist properties of variational posteriors. Two of those papers are cited by the authors as references [2] and [56]. Another one I would add to this list is “$\alpha$-VARIATIONAL INFERENCE WITH STATISTICAL GUARANTEES” by Yang, Pati, Bhattacharya. I wonder if the authors would explain whether the results of those papers could be applied to the variational GP approximations for linear inverse problems, as considered in the authors’ submitted work. If they could, then I would like the authors to give reasons why they decided to study this problem directly rather than applying the theory provided previously in those three papers. If they couldn’t, I would like the authors to briefly describe why. I also have a few other questions and suggestions. In Definition 1 the authors say that a problem is mildly ill-posed if $\kappa_j$ has exponential decay. Do the authors mean polynomial decay? Also in Theorem 1, only at the end of the statement do the authors define what it means for a posterior (or a variational posterior) to contract at a certain rate. I would include a definition of what it means to contract at a certain rate before stating Theorem 1. Moreover, I’m not sure if the notation $S_{++}^m$ used at the bottom of page 4 and the notation $\Pi_u$ used at the top of page 5 are introduced anywhere in the paper before they are used.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

A limitation of this work concerns its practical usability. For example, the authors’ results hold under the assumption that the model is well-specified. It is hard to expect that such an assumption will be satisfied in real-life applications and for real data sets. I believe that the authors should acknowledge this limitation and perhaps comment on possible extensions to instances of model misspecification.

Reviewer Sbdt2023-08-13

Thank you for the explanation. I will keep the evaluation as it is.

Reviewer TgVX2023-08-14

Reply to authors

I thank the authors for their reply which I have read.

Reviewer rih92023-08-16

Reply to authors

The authors have addressed all my comments and questions. I will keep my evaluation as it is. I must say, however, that having read the comment of reviewer VbCA I can't help but agree with them that such highly technical papers might be better suited for publication in Statistics/ML journals. It is difficult to check complicated mathematical arguments carefully within the very tight schedule of the NeurIPS reviewing process. The authors' results look correct but I admit I did not manage to check their proofs carefully and from what I see nor did some of the other reviewers.

Reviewer KBMK2023-08-19

The authors' reply has addressed my previous concerns, and I do see there are slightly more quantitative evaluations in the supplemental material. I've updated my score to 6.

Authorsrebuttal2023-08-21

We would like to thank the Referee for the additional point.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC