Thank you for your thorough review. We address your comments as follows.
**W1. The paper misses some very important references on phase transitions in generative diffusion.** In the revised version of the paper, we now refer and discuss the relations with the work of Raya and Ambrogioni 2023, Ambrogioni 2023, and Li and Chen 2024. Please let us know if this addresses your concern.
**Q1. Could you connect your result to the theory of spontaneous symmetry breaking in generative diffusion models?** Our results can be stated as follows in terms of symmetry breaking. (1) The results from Biroli et al (2024) translated to flow-based generative models show that for generating samples from the two-mode GM in dimension $d$, there is a spontaneous symmetry breaking at times of order $1/\sqrt{d}.$ (2) We then show in Proposition 1 that the critical window where the symmetry breaking happens can be made of constant length in $d$ through a time dilation, which we provide in Eq. (10) in the original paper. (3) Finally, analyzing the learning problem using this time dilation shows that the generative model can resolve the spontaneous symmetry breaking, stated formally in Results 1 and 2, and Corollary 5.
**W2/Q2. The time dilation formula in Eq.10 can only be calibrated on a single symmetry-breaking point. ... It would be more useful to have a formula that can recalibrate the sampling of data with multiple decision points.** The time dilation formula in Eq. 10 can be extended to more than two modes as follows. Consider $\mu=\sum_{i=1}^{m}p_{i}\mathcal{N}\left(r_{i},\sigma^{2}\text{I}\right)$ where $r_{i}\in\mathbb{R}^{d}$ and $|r_{i}|$ goes to infinity with $d,$ but $m,p_{i},\sigma^{2}$ are constant with respect to $d.$ If $X_{t}$ is the generative model associated with the interpolant $I_{t}=(1-t)z+ta$ where $z\sim\mathcal{N}(0,\text{I})$ and $a\sim\mu,$ then $X_{t}$ estimates $p_{i}$ at times of the order $1/|r_{i}|.$ We show this in Proposition 2 in the Appendix, Section E, of the revised version of the paper by arguing that it is only at times of order $1/|r_{i}|$ that the denoiser associated to $r_{i}\cdot X_{t}/|r_{i}|$ is nontrivial. Hence, to estimate $p_{i}$ we require a time dilation $\tau_{t}$ such that there exists $a$ and $b$ with $b-a=\Theta_{d}(1)$ where
$$
\tau_{t} =\Theta_{d}\left(\frac{1}{|r_{i}|}\right) \text{ for }t\in[a,b].
$$
From here we can derive the more general dilation formula that you asked for. We need to specify a dilation that for every $i$ ensures that the condition on $\tau_t$ is fulfilled. Assume $|r_{1}|\leq|r_{2}|\leq\cdots\leq|r_{m}|$, let $n=m+1$ and let $\kappa>0.$ Then
$$
\tau_{t}=\begin{cases}
\frac{\kappa nt}{|r_{m}|} & \text{if }t\in[0,1/n]\\\\
\frac{\kappa(nt-1)}{|r_{m-1}|}+\frac{\kappa}{|r_{m}|} & \text{if }t\in[1/n,2/n]\\\\
\cdots\\\\
\frac{\kappa(nt-(m-1))}{|r_{1}|}+\kappa\left(\frac{1}{|r_{2}|}+\cdots+\frac{1}{|r_{m}|}\right) & \text{if }t\in[(m-1)/n,m/n]\\\\
\left(1-\kappa\left(\frac{1}{|r_{1}|}+\cdots+\frac{1}{|r_{m}|}\right)\right)t+\kappa\left(\frac{1}{|r_{1}|}+\cdots+\frac{1}{|r_{m}|}\right) & \text{if }t\in[m/n,1]
\end{cases}
$$
Then we have that $p_{i}$ is learnt when $t\in[(m-i)/n,(m-i+1)/n]$ and the $\sigma^{2}$ will be learnt when $t\in[m/n,1],$ giving rise to $m+1$ different phases. In the special case of $|r_{i}|=|r_{i+1}|,$ both $p_{i}$ and $p_{i+1}$ will already be learnt in $[(m-i)/n,(m-i+1)/n]$ so that the phase on the interval $[(m-i+1)/n,(m-i+2)/n]$ is unnecessary. Taking this consideration into account when using the general formula we have derived applied for the two-mode GM gives the time dilation formula from Eq. (10) in our paper. The only difference is that the time dilation here maps $[0,1]$ to $[0,1]$ and in the paper we map $[0,1]$ to $[0,2].$ We included this in the Appendix, Section E, of the revised version of the paper. Note that the idea to detect transitions via comparing the coefficient of data and noise, expressed formally in Lemma 7, is quite general and is a way to extend the time dilation to even more general data distributions.