1. For Theorem 3.1, we further clarify its connection to prior works (in Remark 3.3). Based on the referee’s suggestion, we add a sentence immediately before the statement of Theorem 3.1 to demonstrate that the associated time threshold is sharp (Theorem 3.4), which is new.
Regarding Theorem 3.9 in the previous version, we have followed the referee's advice by presenting it, in the revised version, as a standard result (16) rather than formulating it as a theorem. As discussed in our paper, the key focus is on whether these standard estimates can be further improved. We also refined the statements following (16) to emphasize two key points: (1) the $1/t$ gradient bound is generally sharp and cannot be improved, and (2) the Hessian bound can be improved from $1/t^2$ to $1/t$ for ``most" $x$. For the reader's convenience, we kept the proof of (16), as it is straightforward and some steps, such as (32) in the previous version, are used for the proof of Theorem 3.10 (Theorem 3.9 in the updated version).
Regarding Theorem 4.1, we noted in Remark 4.2 in the previous version that it is a special case of more general results (e.g., Chapter IV of Ikeda $\&$ Watanabe (2014)). Since uniform Lipschitz continuity is often assumed as a standard condition for long time existence of solutions for convenience, we believe it would be beneficial for readers, especially non-experts, to explicitly state this special case as a theorem and provide its proof, where uniform Lipschitz continuity may not hold. Thanks to the referee's suggestion, we have added several sentences just before the statement of Theorem 4.1 and in Remark 4.2 to emphasize that this is a special case of known results.
2. We apologize for the inconvenience during your reading, and have revised to list the assumptions in Section 2.2. Regarding whether those technical assumptions are reasonable, for smooth initial data, Assumption 2.3 represents a typical (though not strictly necessary) condition to ensure the existence of solutions. Also, if we assume $D^2g$ is bounded, $g$ grows at most quadratically and $Dg$ grows at most linearly. Hence it is reasonable to expect that $\alpha_1$ behaves similar to $|x_0|^2\sim O(n)$ and $\nabla g(x_0)$ behaves similar to $|x_0|\sim O(\sqrt{n})$. From the computational point of new, the main issue is to track dependence on the dimension $n$ that could be potentially very large.
3. The bound on line 1320 of the previous version is due to (14) of Theorem 3.5 in the revised version ( the same as (11) of Theorem 3.5 in the previous version) and the fact that $\max(a,b)\leq a+b$ after absorbing all other constants depending on $C_0$, $\alpha_2$ and $L$ but independent of $n$. We added this explanation in the revised version.
4. As mentioned in the paper, if the Hessian bound improvement ${1/t^2}\to {1/t}$ can be established for every $x$, then an improved convergence rate of the Wasserstein bound follows immediately from Theorem 3 of Bortoli (2022). Our result shows that this Hessian bound improvement typically holds except on a set of measure zero of $x$. Accordingly, the theoretical assumption of $O(1/t)$ Hessian bound is meaningful in most situations when addressing the convergence rate issue. We share the referee's view that it would be ideal if we could quantify rigorously how this small exceptional set (usually inevitable for non-convex data set) could impact the convergence rate. However, this is mathematically very challenging due to the complex topological structure of the exceptional set, which we plan to investigate in the future. A sentence is added in the revised version to clarify this point.