We would like to extend our sincere gratitude for taking your time reading the note and encouraging timely response from reviewers. We thank all reviewers again for their constructive feedback.
In particular, we appreciate reviewer 1EPQ's efforts in re-examining our technical details. However, we are writing this letter because we feel compelled to express our concern about their comments as they challenge our work either **without careful scrutiny** (the first time) or **with strong prejudice against our problem setup** (the second time), which could have a negative impact on their evaluations.
Moreover, reviewer 1EPQ thinks our presentation were *"sloppy"* (not even fair) while all the other four reviewers rated this particular dimension as good or excellent, which inevitably makes us slightly question their expertise in this field.
We provide specific arguments in response to their feedback as follows:
Since they mentioned *"practical"* issues multiple times, we would like to point out that **practical algorithms are primarily proposed to better solve a problem**. Our practical contribution in this work is the **first deterministic particle-based algorithm** for Gaussian Variational Inference (GVI). Obtaining a complete theory could take years. In this current work, we have already provided appealing theoretical results for the special case of Gaussian targets (even in the discrete-time and finite-particle setting), along with empirical evidence showing superior performance for general targets. While a complete theory would be ideal, it is unfair to judge our algorithmic contribution as *what they called "useless"* simply for the reason that there is currently no complete general bound. Note that even after more than 7 years since the publication of the SVGD paper by Liu \& Wang, there is no complete theory on understanding it. The work of Shi \& Mackey 2022 might be one of the best results so far but is still far from alignment with practice.
Also, reviewer 1EPQ seemed to have strong objections against **bilinear kernels**. As we had made clear in the rebuttal, in Gaussian-SVGD we use bilinear kernels to perform GVI. BWGD (Lambert et al) re-emerged as one of the algorithms under our (bilinear kernel) framework and FB-GVI (Diao et al) improved upon BWGD. If the bilinear kernel were really impractical as the reviewer had claimed, these two previous methods should have been regarded as impractical as well. And by similar logic, one can even argue that the whole field of GVI were impractical. Furthermore, regarding FB-GVI, the reviewer's claim that **"high-probability guarantees should follow easily"** is not fully correct. It would need strong assumptions on the stochastic gradients to derive a high-probability bound whereas our results are entirely **deterministic**. This further highlights the advantage in having an algorithm with deterministic particle updates for GVI, complementing the stochastic algorithms like BWGD and FB-GVI.
**Regarding Theorem 3.7**, as we had pointed out in the rebuttal, even for the continuous time setting, we provide the **first uniform-in-time** propagation of chaos result. It is unfair to say that there is lack of contribution due to the continuity of time. Apparently **uniform vs non-uniform** is more significant than **continuous vs discrete**, and getting a continuous-time result is always a first step towards understanding discrete-time dymanics. **Last but not least**, the idea of "uniform-in-time" comes from the general observation that the difference of some summary statistics between two particle systems (governed by the same ODE) with different initializations reduces after some bounded time rather than keeps growing exponentially in time. Although our current proof relies on the precise dynamics, the idea is generalizable (e.g. from one particular ODE to an ODE class) and **definitely NOT an overclaim**.
There are other minor points from their response that do not sound very reasonable to us. For example, they claimed that sampling Gaussian with known $b$ and $Q$ is *"unrealistic"*, and carefully clarified that $b$ and $Q$ are indeed known in our paper. However, the usual setup of sampling or GVI is exactly that **the target density is known up to a normalization constant** (e.g. see the SVGD paper by Liu \& Wang).
In summary, we **have serious doubts about the reviewing principles of reviewer 1EPQ**. On one hand, they *sentenced (please forgive us for using this word due to their sarcastic tone)* our algorithm as impractical only based on the theoretical setting rather than experiments, which is **not a justified way to evaluate algorithms**. On the other hand, while our theoretical results aim to provide new insights to SVGD and GVI, they seemed to believe that theoretical insights are *"useless"* unless they immediately buy you results for general settings.
We sincerely hope that you could take our concerns into consideration and make a fair judgment!