Thank you for your positive review and thoughtful questions. Please find our responses below.
*The paper poses a question on SVGD, but then in fact does not answer this question (What does SVGD, $\lambda=0$, converge to?), nor does the bulk of the paper seem particularly interested in answering it.*
When $ \lambda = 0 $, the algorithm converges to a set of measures that includes the target $ \pi $ as well as some probability measures without density (the addition of the noise allows to avoid measure witout density). However, you are correct that we should not have emphasized this point. As you suggested, we have removed the italics in the updated version.
Indeed, a large noise ($\lambda \gg 1$) will not bias the algorithm. As demonstrated in the newly added Section C, by setting $\lambda$ large, NSVGD behaves like the Langevin algorithm. In fact, it effectively becomes a Langevin algorithm with a step size of $\lambda\gamma\_k$. Therefore, to ensure convergence, it is necessary to have $\lambda\gamma\_k \ll 1$.
*Do the theorems allow any uniformity in the dimension d of the underlying problem, say in the case of a log-sobolev constants?*
In our proof technique, we cannot establish bounds that provide a definitive answer. However, we can state that the number of particles ($n$) should depend exponentially on the dimension. Indeed, the larger the space, the more particles are needed to adequately cover it. Moreover, the time has a linear dependence on the dimension. Specifically, by examining the proof of Lemma 7, we observe that the constant $C$ (which serves as an upper bound for the Kullback-Leibler divergence) depends linearly on $d$. Additionally, the log-Sobolev constant corresponds to the convexity constant of $F$, where we recall that the target distribution is proportional to $\exp(-F)$.
To summarize, we believe there exists a strong dependency on the dimension $d$.
This version corrects grammatical issues and improves the clarity and flow of the text.
*Can you provide further experiments, evaluating performance of NSVGD vs SVGD and Langevin sampling? Is there any setup, in particular where lambda=0 is favorable?*
We hope to satisfy your requirements in Section C of the updated version of our manuscript. To summarize, we used the toy example of a Neal funnel density. In this setup, SVGD outperforms the Langevin algorithm. We show that by setting $\lambda$ small enough, NSVGD retains the performance of SVGD. Therefore, NSVGD is advantageous as it performs similarly to SVGD and better than Langevin, while also providing convergence guarantees. But, in this particular setup, $\lambda=0$ yields better performance.
*The law of those $n$ points may very well have a density, and moreover the law of a randomly selected point could be close to $\pi$. Why is the discreteness of this measure a problem? (You also mention this on l84). Furthermore, even if this were a problem, shouldn't one just change the metric?*
By stating that SVGD converges to $ \pi $, we mentioned in line 51 that this means the empirical measure $\mu\_k^n $ converges to $\pi $. In this case, for SVGD, one can only hope for
$ n \to \infty $.
As you said, one could alternatively define the convergence of SVGD as the law of a single particle converging to $ \pi $.
But, SVGD works because, in the population limit $ n = \infty $, the Kullback–Leibler divergence $ D\_{KL}(\mu\_k^\infty \| \pi) $ decreases along $ k $. This property cannot hold for finite $ n $, since $ D\_{KL}(\mu \| \pi) $ is defined only for measures $ \mu $ admitting a density. To the best of our knowledge, $ D\_{KL}(\text{Law}(X\_k^{1,n}) \| \pi) $ does not decrease for $n<\infty$.
*Why should all particles need converge, instead of the evolution of a tagged particle converges?*
By Assumption 1-(ii), the particles at time $k =0$ share the same distribution. Furthermore, by the definition of the algorithm, we can verify that all particles $X\_k^{1,n}, \dots, X\_k^{n,n}$ also share the same distribution at any time $k$. Consequently, we can only expect all the particles to converge or none at all.
*l.386. There is some double limit notation not properly introduced, l.517. ``usefull''*
Thank you for your careful reading. This has now been corrected in the updated version (the definition of the double limit is in l.139.).
**We agree with you that it would be interesting to establish a convergence result when the dimension $ d $ grows with $ k $ and $ n $. However, at present, we are unable to establish such a result. We hope we have clarified the reason why it is necessary to study the set of all particles rather than focusing on a single selected particle. Additionally, we hope the new Section C in the updated manuscript provides clarity on the behavior of NSVGD compared to SVGD. If you are satisfied with our response, we would be deeply grateful if you would consider raising your score.**