Response to Reviewer Exbp
We thank the reviewer Exbp for their attentive reading: we will correct all typos and better structure the related works paragraphs.
We respond below to their other questions and comments.
**Applicability of Corollary 5 ---** the corollary requires initializing Langevin dynamics at the proposal, i.e. $p_0 = \nu$,
but we do not require that $\lambda_0 = 0$.
We refer the reviewer to the short proof in Appendix C.1:
to control $\mathrm{KL}(p_0\|\mu_0)$ we use the tempering rule,
and ultimately simply bound $\lambda_0$ by $1$.
**Theorem 9 in case $\lambda_i$ is close to $1$---**
We are in complete agreement with the reviewers' intution,
and note that in Theorem 9, $\delta_k$ is defined as $\delta_k := 8m^2 e^{-(1 - \lambda_k)m^2}$,
so that when $\lambda_k$ is close to $1$, there is no exponential
slow-down. Actually, Theorem 9 makes this intuition quantitative
and states that so long as $1 - \lambda_k \leqslant C m^{-2}$,
the slow-down will only be of polynomial order in the mean separation.
**Choosing an optimal proposal distribution knowing some information about the target distribution ---** consider the setup when the target is a mixture of two symmetric Gaussian distributions and we initialize with a Gaussian located “in between” the two target modes, specifically at the barycenter. Then, convergence is provably fast for Tempered Langevin [1, Example 1] and related sampling processes [2]. Yet, initializing in this way requires knowing the locations of all target modes: this would require solving a “global optimization” problem, which can be a harder problem than the original problem of sampling from the target distribution [3].
**Do the lower bounds hold in higher dimensions? ---**
We expect that the lower bounds go through in higher dimensions
with minor modifications. The key high-level point in all
of the lower bounds is that at intermediate times,
the tempering path becomes bimodal, and if the
law $p_t$ of the particle following the tempering
is too concentrated in one of the modes vs. another,
it will take exponential time to spread mass between the modes
(for example, see Prop. 22 in Appendix D).
In other words, and as illustrated in Fig. 4 Appendix E, the core phenomenon behind our lower bounds is the fact that along the geometric path, mass tends to “teleport” from one mode to another, preventing Langevin to converge as the particles eventually get stuck in the first encountered mode.
All of this translates intuitively, and likely rigorously,
without issue to higher dimensions.
We chose to present the lower bounds in dimension $1$
for simplicity, as well as
to make it clear that the problem is not a curse of dimensionality (which may be common across methods),
but instead a problem specifically with the tempered Langevin itself.
**Why additional dissipativity assumption? ---**
Our only use of the dissipativity assumptions
is to control the second moment of $p_t$, the law of the process $X_t$ given in (9).
At a technical level, the reason why we have this extra assumption
as compared to the standard analyses of vanilla Langevin
is precisely because of the additional terms arising from
the tempering. Weakening this assumption further
is an interesting direction for future work.
**Convergence under weaker functional inequalities
than log-Sobolev? ---**
The main technical novelty of our analysis is the
way that we deal with the new terms
arising from the tempering dynamics (see Step 1 and Step 2 in Appendix A.2, as well as the supporting Lemmas in Appendix A.4).
In particular,
the extra terms which arise here are particularly suitable
to analysis when the Lyapunov function is $\mathrm{KL}$.
For example, we are not
aware of a straightforward means of extending
our analysis to $\chi^2$.
Since these alternative functional inequalities imply
convergence in alternative Lyapunov functions (e.g. Poincaré involves
analysis in $\chi^2$), we are therefore not aware of a straightforward
extension of our results to weaker isoperimetric assumptions.
[1] Guo et al. Provable Benefit of Annealed Langevin Monte Carlo for Non-log-concave Sampling. Arxiv, 2024.
[2] Madras and Zheng. On the Swapping Algorithm.
Journal of Random Struct. Algorithms, 2003.
[3] Ma et al. Sampling can be faster than optimization. PNAS, 2019.