Gradient-Free Kernel Stein Discrepancy

Stein discrepancies have emerged as a powerful statistical tool, being applied to fundamental statistical problems including parameter inference, goodness-of-fit testing, and sampling. The canonical Stein discrepancies require the derivatives of a statistical model to be computed, and in return provide theoretical guarantees of convergence detection and control. However, for complex statistical models, the stable numerical computation of derivatives can require bespoke algorithmic development and render Stein discrepancies impractical. This paper focuses on posterior approximation using Stein discrepancies, and introduces a collection of non-canonical Stein discrepancies that are gradient free, meaning that derivatives of the statistical model are not required. Sufficient conditions for convergence detection and control are established, and applications to sampling and variational inference are presented.

Paper

References (50)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer WN2C4/10 · confidence 3/52023-06-19

Summary

The paper provides a method to estimate the kernel Stein discrepancy without gradients.

Strengths

The paper provides a method to estimate the kernel Stein discrepancy without gradients.

Weaknesses

The paper is not clearly written for someone who is not familiar with the field. After reading, I'm still confused about whether someone else has published a similar method before.

Questions

How limited is the class of distributions that allow gradient free estimation of the kernel Stein discrepancy? And the paper probably cherry picked good experiments, how often do the mentioned failure modes occur?

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

The paper is not clearly written for someone who is not familiar with the field. After reading, I'm still confused about whether someone else has published a similar method before.

Reviewer DoGh8/10 · confidence 3/52023-06-21

Summary

This paper explores the use of Stein discrepancies in scenarios where the score function of the target distribution is unavailable or computationally impractical to evaluate. The authors introduce a novel approach called gradient-free kernelized Stein discrepancy (GF-KSD), which leverages a Stein operator developed by Han and Liu (2018) that does not rely on the score function. The authors establish sufficient conditions for the resulting divergence to control and detect convergence. The empirical evaluation of this divergence extends its application to Stein importance sampling and Stein variational inference, surpassing the scope of the previous work by Han and Liu (2018).

Strengths

**Theories**: Whilst the gradient-free Stein operator has been previously discussed in the community (Han and Liu, 2018), its theoretical properties have not received sufficient attention. This paper established conditions under which the GF-KSD can detect and control convergence of a sequence of empirical distributions to a target probability measure (Theorem 1 and 2). To my knowledge, this contribution fills an important gap in the existing literature and significantly enhances our understanding of the topic. **Discussions**: Key results on both the theoretical and empirical sides are sufficiently discussed. Limitations of the GF-KSD are thoroughly examined and supported by empirical evidence (Section 3.3). Discussions on the choice of the degree of freedom, the density, are also included (Section 3.1). **Structure**: This paper exhibits excellent writing with clear motivations throughout. J. Han and Q. Liu. Stein variational gradient descent without gradient. In Proceedings of the 35th International Conference on Machine Learning, pages 1900–1908. PMLR, 2018.

Weaknesses

**Experiments**: My only major concern is over the empirical results. The problems examined in the experiments are primarily toy examples with relatively low dimensionalities. where the one with the highest dimension is a 8-dimensional inference problem for a Lotka-Volterra model. The highest-dimensional problem considered is an 8-dimensional inference problem for a Lotka-Volterra model. To further validate the applicability of the proposed method in more realistic scenarios, additional empirical evidence on higher-dimensional problems would be beneficial. Specifically, it would be interesting to investigate the performance of the default choice of Laplace approximation when the dimensionality of the target distribution is high. This analysis could shed light on whether the default approach continues to yield satisfactory results in non-toy scenarios. Moreover, the reported numerical instability issue in Section 4.2 warrants further investigation, particularly in high-dimensional problems. Understanding the extent to which this instability becomes prominent in higher dimensions is crucial as it may impact the practical viability of the GF-KSD approach.

Questions

It would be helpful to elaborate on the numerical instability issue noted in Section 4.2. Does this issue only occur when using GF-KSD for Stein variational inference, or does it also occur in the experiments in Section 4.1? Is that an artefact of the specific form of the target densities chosen for this experiment? Can this be avoided by a judicious choice of $q$?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

Examples where the proposed approach can fail are extensively discussed in Section 3.3.

Reviewer 9g9f4/10 · confidence 4/52023-07-05

Summary

This paper proposed a posterior approximation using a new Stein discrepancy, which does not require derivatives of the statistical model. For that purpose, the authors derived the new discrepancy, called gradient-free KSD, and studied its statistical and convergence behaviors theoretically. Then the authors developed algorithms for differential equations, for which stable computation of derivatives is difficult, and a new sampling algorithm that bypass the Hessian calculations.

Strengths

- A new KSD that does not require gradients, similar to the idea of importance sampling, is proposed. This leads to new algorithms for differential equations where the gradient is difficult to compute and for problems requiring Hessian calculation. - The authors studied the theoretical property of the proposed GF-KSD by extending the existing KSD theory. - Not only theoretical analysis but also detailed numerical investigations of the choice of parameters and $q$ are carried out with the actual use in mind.

Weaknesses

- The writing style is such that the main paper alone is not complete, and it is assumed that the reader will read the Appendix. For example, Eq. 6 of Line 115 does not appear in the main text, and the tilted Wasserstein distance defined in Theorem 1 is introduced without any explanation of its properties in the main text. - The writing style could be improved since the discussion about existing research and the explanation of the proposed method are mixed, making the paper difficult to read. -Some parts are mathematically undefined or under-discussed - In Definition 2, sup is undefined - In Line 158, at last, $\not \to$ is undefined. - I don't know how widely the tilted Wasserstein distance (TWD) in Theorem 1 is known to the general public, but there is no discussion of the properties of TWD. Therefore I could not understand how important Theorem 1 is, that is, how important it is when it is said that convergence of TWD leads to convergence of GF-KSD; even after reading the proof of Theorem 1, I could only understand that TWD is a convenient form of the usual Wasserstein distance, which is obtained after applying the triangle inequality. - I could not understand the importance of the proposed method because I am not sure for what problems the proposed method is effective. I agree that it may be useful for differential equation problems, but the authors only applied the method to very small models of Lotka-Volterra in the experiments. In such a setting, MCMC is the standard approach, and even if the model is high-dimensional, we can solve it efficiently by variational inference. Also, although the combination with Stein variational inference seems interesting, I wondered if GF-KSD is really flexible enough to generate samples for complex real data under the two restrictions suggested in Section 3. One restriction is that the tail of $q$ should not be far from the target distribution; the other is that it must not be high-dimensional. I think the application to differential equations seems promising, so it would be better to find a problem setting where GF-KSD is more useful than MCMC and standard variational inference.

Questions

I would appreciate it if the authors would answer the concerns described in Weakness. As for minor questions; - What do the dotted and solid lines in Figure 1(b) correspond to? - Is it required to adjust parameters of the Laplace distribution, KDE, and GMM for $q$ in some way ? If so, what is the recommended method ? - Looking at Figure 3 (b), it seems that the number of samples ($n$) must be very large ($\log n=5$, i.e., $n=150$) even for low-dimensional problems such as d=8 in order for there to be any difference in energy distance. Is my understanding correct?

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

The limitation of the proposed method is discussed in detail.

Reviewer SmKF6/10 · confidence 3/52023-08-07

Summary

The authors study the kernel Stein discrepancy based on a Stein operator, introduced previously in [Liu 2018], which does not require access to the gradient of the target density. As their main theoretical result, they prove that, under certain assumptions on the target distribution (and the auxiliary distribution q) and the approximating sequence, the discrepancy controls weak convergence. The authors provide recommendations for the choice of the auxiliary distribution q in the construction of the discrepancy and present an experimental study which identifies certain failure modes of the discrepancy. They also discuss two applications to posterior approximation.

Strengths

The idea of the authors to use the gradient-free Stein operator in order to define a new kind of KSD is interesting and worth studying. There are certainly a lot of situations where the evaluation of the gradient of the model is costly in practice. Giving the practitioners an option to avoid evaluating it when computing the KSD is certainly useful. The discussion of the properties of the new GF-KSD is very thorough and based on both theoretical results and a substantial experimental study. I particularly liked the the experiments revealing failure modes of GF-KSD as those are not necessarily captured by the theoretical results.

Weaknesses

The main theoretical contribution of the paper seems to be provided by Theorem 2. However, I am a bit concerned about its real applicability. The authors require that the auxiliary distribution $q$ have tails heavier than the target $p$. On the other hand, they require that $q$ is distantly dissipative, which precludes the use of heavy-tailed $q$. This all means that $p$ must not be heavy-tailed if the assumptions of Theorem 2 are to be satisfied. Moreover, the authors suggest four potential choices of $q$ (Prior, Laplace, GMM and KDE) but note that GMM and KDE are impractical in general. At the same time, the Prior method does not satisfy the assumptions of Theorem 2 for heavy-tailed priors. The Laplace method, on the other hand, seems to satisfy the assumptions of Theorem 2 only for targets which are sub-Gaussian.

Questions

Related to what I wrote above, could the authors state clearly what class of targets $p$ their Theorem 2 applies to? Could they also try to characterise the class of targets $p$ for which one can easily construct a useful auxiliary distribution $q$ in practice (using one of the methods of section 3.1), such that $q$ and $p$ satisfy the assumptions of Theorem 2?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

I believe the real applicabilty of the theoretical results presented by the authors is somewhat limited, to an extent that is not clearly acknowledged in the paper. But I am looking forward to reading the authors' response as perhaps I'm missing something.

Reviewer DoGh2023-08-12

Thank you for your response. With that said, I still think a non-toy numerical example would significantly improve the paper -- the proposed method is advertised to be a rescue when the score function of the statistical model is impractical to evaluate; despite the fact that GF-KSD is more practical than the standard KSD in these cases, for the presented examples many alternative methods exist and have been demonstrated to work reasonably well, making it unclear whether this method is practically useful (see also the comments by Reviewer 9g9f). Also, my concern of whether the default choice of Laplacian approximation can still give good performance in non-toy, high dimensional examples is not yet answered. Due to the above, I have kept my scores, but happy to re-evaluate if the authors provided convincing numerical evidence to address these concerns before the rebuttal period ends.

Reviewer SmKF2023-08-16

Reply to the authors

Thank you very much for your answers. I have increased my rating of the paper.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC