Detailed answers to the questions (part 2)
5. We thank the referee for pointing out that we should clarify further the panel E of Fig. 4. The aim of this figure was to show the “hysteresis” phenomenon which is a clear signature, devised and observed in statistical physics, to unveil that a high-dimensional probability measure had a phase transition at which it split in two different lumps. The procedure consists in tilting the probability measure by introducing in the energy function a contribution that favors one lump over the other. If the measure is indeed concentrated on two distinct lumps, one induces a sudden transition (“first-order transition” in physics). In our case, the lumps are associated to the learned patterns and this extra contribution consists in the scalar product between the visible variables and the learned patterns times $h$ (the field controlling the strength of the tilting). In the presence of a first-order phase transition, one generically finds the phenomenon of hysteresis, i.e. the transition from one lump to the other can be retarded because of metastability, thus leading to the characteristic hysteresis loops that we indeed show in Fig.4 E (see e.g. den Hollander, Frank. "Metastability under stochastic dynamics." Stochastic Processes and their Applications 114.1 (2004): 1-26; Bovier, Anton. "Metastability." Methods of contemporary mathematical statistical physics 1970 (1970): 177-221 for rigorous treatment and Chaikin, Paul M., Tom C. Lubensky, and Thomas A. Witten. Principles of condensed matter physics. Vol. 10. Cambridge: Cambridge university press, 1995 for a physics treatment). This figure therefore gives direct evidence of the decomposition of the measure in distinct lumps corresponding to the learned patterns, and that this decomposition takes place at the second-order phase transition happening during learning. We will explain better these points in the revised version.
6. To start with ref[8], this work presents a mean-field theory describing theoretically the static behavior of RBMs assuming a proper analytical form for the weight matrix. In this work, the phase transition for such a setting is derived but only the static behavior is analyzed theoretically. Our present work instead characterizes the dynamical behavior of the RBM. It shows at which rate the weight matrix is shaped by the dynamics, it describes the different mechanisms (evolution of the weights into a preferred direction, by which time evolution, etc.) in details, and we further show that this is what happen when performing a training in this regime. Also, we show how the weights evolve in the case where the clusters of the dataset are correlated.
Concerning the relationship between sec 4 and 5, the reviewer’s comments helped us to understand that we have to clarify the relationship between the SVD decomposition used in Sec. 5 with the projections used in Sec. 4. There is a direct connection between the PCA and the preferred directions ($\xi^1 + \xi^2$ and $\xi^1 -\xi^2$) used in the theoretical analysis, but it was not clear enough in the current version. You could see it from the appendix, but the point was not made explicitly anywhere. We will discuss this and make it explicit in the revised version. In fact, by looking in the appendix at the eq. after line 491, we decomposed in our model the correlation matrix on the vectors $\eta$ and $\bar{\eta}$: this is the SVD decomposition for this very simple model. It is therefore clear that since the model exhibits a growth first towards \eta, the first eigenvectors of $\langle s_i s_j \rangle$ and then in the direction of $\bar{\eta}$, the second direction, we are indeed proving that the dynamics is driven by the PCA of the dataset.
We think that the part on the divergence of susceptibility is the clearest link between the two sections. Section 4 shows that the analytical behavior is confirmed in numerical experiments with real data sets: first, the relation between the PCA and the SVD of the RBM weight matrix, second, the phase transition: how the susceptibility diverges (as the system size increases and with a known exponent) as well as the divergence of the mixing time.
We will make clearer in the final version the link between the two parts.