From Biased to Unbiased Dynamics: An Infinitesimal Generator Approach

We investigate learning the eigenfunctions of evolution operators for time-reversal invariant stochastic processes, a prime example being the Langevin equation used in molecular dynamics. Many physical or chemical processes described by this equation involve transitions between metastable states separated by high potential barriers that can hardly be crossed during a simulation. To overcome this bottleneck, data are collected via biased simulations that explore the state space more rapidly. We propose a framework for learning from biased simulations rooted in the infinitesimal generator of the process and the associated resolvent operator. We contrast our approach to more common ones based on the transfer operator, showing that it can provably learn the spectral properties of the unbiased system from biased data. In experiments, we highlight the advantages of our method over transfer operator approaches and recent developments based on generator learning, demonstrating its effectiveness in estimating eigenfunctions and eigenvalues. Importantly, we show that even with datasets containing only a few relevant transitions due to sub-optimal biasing, our approach recovers relevant information about the transition mechanism.

Paper

References (50)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer BQ7h8/10 · confidence 3/52024-07-10

Summary

@Authors, I would appreciate it if you could point out any inaccuracies in the following summary since it took me a long time to understand your paper and am still not completely certain my understanding is correct. The paper aims to address a common problem for molecular dynamics simulations, which is to obtain collective variables that can be used to bias simulations such that, e.g., rare transitions occur more frequently and important observables can be computed faster. The authors follow previous work and learn the system's generator, whose eigenvalues inform transition timescales and whose eigenfunctions can be used as collective variables to bias simulations along. The novelty is in deriving an approach that allows learning the generator from biased enhanced sampling simulations which converge faster and observe more transitions than conventional simulation. The authors evaluate their method on toy systems.

Strengths

1. Novel approach to accelerate obtaining scientifically important collective variables by enabling their computation from simulations biased by enhanced sampling. This approach requires developing new theory and to derive a training objective. The procedure to do so seems non trivial to me but I cannot sufficiently judge the value of the provided theory as I can only follow the broad strokes and am not well versed in the used math. I hope other reviewers can provide more useful signal for this aspect. 2. Impactful application: The paper provides useful ideas for a problem with significance for scientific applications that can have large downstream impacts. 3. Very well written introduction and overview of the transfer operator and generator formalism.

Weaknesses

My points contain many understanding questions and some "criticizing" questions. It would be great if you could also answer the understanding questions. 1. Experiments: The experiments are carried out on three toy systems and the proposed approach visually improves upon two deep learning baselines in one of the three experiments. 1. I understand that the cited works on generator learning provide the same or fewer experiments. However, from skimming the cited works it also seems to me that one of the main areas of application would be simulations on proteins. Is that incorrect and if not, why do you and others not provide experiments on proteins? Aren't there proteins with well known dynamics on which we could evaluate the methods? 2. **I think this method should be tried on systems of increasing sizes until it fails. The experiments of a paper introducing a new method ideally provide information on the capabilities of a method.** Why do you (and the rest of the generator learning papers) not do so? 3. Why do you omit the deep learning baselines from the ALDP experiments? 2. For an understanding of how your method fits into the broader landscape of approaches to determine CVs, I think it would be important to also provide comparisons with non-deep learning approaches. The relationship to DL methods is clear but to understand the general "usefulness" it would be nice to have delineation from classical methods in terms of approach and experiments. Are there classical approaches that can be used as baselines that would outperform all DL methods? 3. Speed comparisons: it seems to me that the underlying tradeoff for all methods is between speed and accuracy and that accuracy alone might be less meaningful. Is there nothing to be said about the runtimes for data collection and training? Minor: 1. To make the paper more broadly accessible to the wider deep learning community, I think it would be helpful to describe your final operator learning approach more procedurally on e.g. the ALDP example. What is the neural network input and output dimensionality? Why do you have separate neural networks for each operator output dimension? What is the input and output to your neural network? Once you trained your neural network, how do you obtain the eigenfunctions from it. 2. How do you obtain your plots in e.g. the ALDP figure Figure 2.? To understand this: what concretely is the eigenfunction in practice once you computed it? How is it represented? Once you have it, how does it assign different values to each molecular structure?

Questions

I would appreciate any time you can take to answer some of my questions. 1. 107 "for the transfer operator we can only observe a noisy evaluation of the output": Do you mean a noisy version of the expectation that defines a transfer operator? Or that the output one can observe after simulating stochastic dynamics is "noisy"? 2. I assume your underlying goal is to discover CVs as the eigenfunctions of your learned generators. Your generators can be learned from biased enhanced sampling simulations where e.g. more transitions occur. Then we can obtain the CVs from your generator. With the CVs we can run biased simulations. Why are these biased simulations more informative than the biased simulations which you ran to train your generator? Can we compute different quantities from your CV biased simulations? Is it just a matter of your CVs being better biases than the bias of the original enhanced sampling bias?

Rating

8

Confidence

3

Soundness

4

Presentation

2

Contribution

4

Limitations

The paper discusses some limitations such as not being time dependent - the significance of which I cannot assess. Limitations such as the limited evaluations and missing understanding of when the method fails seem more significant to me but are not mentioned or explained why they are present.

Reviewer BQ7h2024-08-12

Thank you very much for confirming some points, and for addressing the reference. I would like to take the freedom to raise my score from 6 to 8 and to increase my confidence.

Authorsrebuttal2024-08-12

Acknowledgement to the reviewer

We are happy that our replies were helpful. We thank reviewer for all their comments and discussions, which I will improve our paper. We commit to incorporate them all in the revision.

Reviewer txkX7/10 · confidence 1/52024-07-12

Summary

In this paper, the authors investigate the possibility of estimating the leading eigenvalues and the corresponding eigenvectors for the evolution operators of Langevin dynamics using biased simulations. To this end, they rely on strong statistical guarantees and on the use of deep learning regression to build a suitable Hiltbert space that approximates the evolution operator of the unbiased process using only biased data. They evaluate the reliability of this approach in three increasingly complex problems and show that this approach successfully recovers the slowest relaxation modes in toy models and obtains better approximate values for the eigenvalues than competing state-of-the-art methods.

Strengths

The work deals with a very important problem in chemistry and physics, namely the identification of the slowest modes of a dynamic process by means of simulations. In complex cases, simulations are not long enough to observe the most prominent transitions, and enhancing sampling methods are used to facilitate the observation of rare events. The problem is that it is often difficult to identify the key collective variables to facilitate the required slow movements. In this paper, the authors propose a simple way to do this using biased simulation data based on strong theoretical guarantees.

Weaknesses

I find it difficult to read the paper and follow it, perhaps because I am too far from the field

Questions

I did not understand well the physical meaning of the modes they extract with their approach. Could they give a physical explanation of them or of the order parameter associated?

Rating

7

Confidence

1

Soundness

3

Presentation

3

Contribution

3

Limitations

Authors discuss correctly the limitations

Reviewer MSBv6/10 · confidence 4/52024-07-13

Summary

This paper studies an infinitesimal generator approach for learning the eigenfunctions of evolution operators for Langevin SDEs. Due to the slow mixing caused by the high potential barriers, direct learning from simulation data can be sample inefficient. Biased simulation (based on a biased potential) is used to explore the space faster, and importance weights are constructed to get an unbiased loss function; such unbiased construction is more natural and feasible in the generator approach compared to the existing transfer operator approach. The minimizer of the resulting quadratic loss function, given a dictionary of basis functions, can be obtained by ridge regression type algorithms as it is a linear method. The authors also propose a loss function to learn a good dictionary that approximates the space of eigenfunctions more accurately; this nonlinear dictionary learning improves accuracy. Experiments on molecular dynamics benchmarks demonstrate the effectiveness of the approach.

Strengths

The introduction and motivation are exceptionally well written and demonstrate the benefits of the generator approach compared to the transfer operator approach for biased simulations. The numerical experiments, especially in Figure 1, show significant improvements in accuracy compared to previous methods.

Weaknesses

The mathematical presentation of the technical details, namely sections 4 and 5, could be made clearer. Many notations are introduced, but the ideas seem simple; see Questions. I would appreciate an overarching description highlighting the key idea and the difference between the proposed approach and existing works using generators.

Questions

- Is the estimator in Section 4 simply the Galerkin approximation of the inverse of $\eta I - \mathcal{L}$ with the basis functions in $\mathcal{H}$? I found the descriptions in Section 4 complicated and not easy to digest. As an example, for equation (11), by definition, the formula will differentiate the term $\|\chi_\eta(x) - \hat{G}^Tz(x)\|_2$ (which is not differentiable at zero) which looks strange. And I didn't understand the sentence "we contrast the resolvent ..." on page 5, line 182: what do the authors mean by "contrast" here? And the explanation of motivation of using the generator regression problem (11) rather than the mean square error also appears less clear. - I am also curious about the mechanism behind the improvement of accuracy in Figure 2. It appears to me that the approach in this paper first uses the loss function in equation (19) to find approximate eigenfunctions and then uses the generator approach in section 4 to refine the estimation of the eigenfunctions and eigenvalues. If so, an understanding of which step of the two plays the significant role will be useful. For example, if the authors apply the generator approach in section 4 to the approximate eigenvalues obtained in the work of Zhang et al. (2022), will the resulting accuracy in Figure 2 also be significantly improved? - Potential typos: Page 2, line 67: "principle" -> "principled". Page 3, line 125-126: the notation of $C_{\gamma}$ is not introduced. Moreover, it seems $\lambda_i = \log \mu_i / t$ rather than $\log (\mu_i/t)$. Page 8, line 282, "approqch" - Page 3, line 129: "If t is chosen too small, the cross-covariance matrices will be too noisy for slowly mixing processes" Could the authors elaborate more on this?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Limitations are discussed.

Reviewer MSBv2024-08-09

Thank you for the detailed response. I have some follow-up points. - **Q1**. I now agree that the estimator is a regularized Galerkin projection of the resolvent onto the space of basis functions, using the energy inner product. This interpretation seems more convenient from an expository standpoint. It would be helpful if the authors can incorporate this perspective into the paper and clarify the writing. In my experience, the current mathematical exposition, with its various operators, norms and formulas, did not provide a very pleasant reading experience. - **Q2**. Yes, it is Figure 1. Thank you. The explanation and results are interesting. Do the authors have any insights regarding the superiority of the loss (19) for representation learning? It seems to me the loss is not that simple and easy to optimize or compute. There is a penalization term to make the $z_i$ normalized and one has the hyper parameter $\alpha$ to tune. Furthermore, based on the authors' response, do I understand correctly that the primary improvement stems from the loss function (19) for representation learning, which is more significant than the regression formulation proposed in the paper? If so I think these insights should be added to the paper.

Authorsrebuttal2024-08-09

Further discussion

We thank the reviewer for the prompt reply. - __Q1__ We agree that Galerkin projection interpretation seems more convenient from an expository standpoint. So, based on the reviewer’s feedback, we propose to modify the content in lines 185-191, by briefly explaining why Galerkin projection should be done in energy space $\cal{H}^\eta_\pi$ instead of $\cal{L}^2_\pi$ space. Then, since the statistical learning risk formulation is the key to understand how to generalizes w.r.t. (unobserved) true distribution $\pi$ from biased simulations, we briefly discuss its equivalence with the Galerkin viewpoint. We believe that this change can help the reader to better grasp the main idea behind the energy, and connect it to the material that follows. Does the reviewer find this proposal adequate? - __Q2: about the loss:__ To answer this question, let us compare (19) to the two losses used in the most related works, [24] and [46]. While the former one (only tested on unbiased dynamics of a toy system) depends on one hyperparameter, the latter loss has $1+m$ hyperparameters, where $m$ is the latent dimension. Importantly, __both losses suffer from statistically biased estimation__ of gradients when optimized over the sample distributions, which may negatively impact the optimization. Moreover, as reported in [46] (see also lines 287-289) additional $m$ hyperparameters in loss of [46] are delicate to tune. __In sharp contrast__, our empirical loss, as stated in Theorem 2, __is an unbiased estimate__ of the true loss in Equation (18), and we didn’t experience any particular difficulty in tuning our only hyperparameter $\alpha>0$. Note also that, in principle $\alpha>0$ is optional to speed up the training, and we may as well use $\alpha=0$. Concerning the computation, we respectfully disagree that the loss is hard to compute. Indeed, for two size $n$ batches of data, it just relies on computation of covariances of features $C$ and their gradients $W$, which is of the order $\cal{O}(n m^2 d)$. Since in the molecular dynamics setting $m$ is typically small, computing the loss is very efficient, similarly to losses in self-supervised learning based on canonical correlations, see e.g. [A]. The only computational bottleneck lies in using the gradients for $W$. However, we believe that this is a necessary price to pay to be able to generalize properly from biased simulations, as motivated in Sec. 3. Finally, we remark that optimization of our loss follows interpretable dynamics, see Figure 3 of the Appendix where plateaus reveal discovery of new relevant features, helping practitioners to decide when to stop the training. - __Q2: about the contribution:__ As the reviewer rightfully recognizes, introducing loss (19) for representation learning is an important contribution of this work. However, our main focus is on __biased simulations__, or, in other words, __learning dynamics that ML algorithm did not observe__. With this in mind, our main contributions are summarized in lines 67-74. For the reviewer's convenience we rephrase and slightly expand them here. To the best of our knowledge, for the first time - __(1)__ we __formalize the problem of learning__ the spectral decomposition of the infinitesimal generator __from biased simulations__ as a regression, - __(2)__ we __derive empirical estimators with guaranteed statistical consistency__ (Theorem 1), and - __(3)__ we propose an end-to-end __representation learning + regression__ pipeline that efficiently scales to large problems. - __(4)__ we __empirically validate our methodology__ on molecular dynamics datasets of increasing complexity. That said, __we will follow reviewer’s suggestion and emphasize better the important role of representation learning__ for the generator regression. Once more we thank the reviewer, and if any other doubt remains or clarification is needed, we are happy to discuss in more detail. Ref. [A] Balestriero et al., A Cookbook of Self-Supervised Learning. Arxiv:2304.12210 (2023).

Reviewer MSBv2024-08-13

Acknowledgement of rebuttal by authors

Thank you for the detailed response. They are very clear.

Reviewer txkX2024-08-09

I really appreciated the authors' explanations, including the new figures (pipeline and the experiment with the proteins). I will change my rating to accept.

Reviewer BQ7h2024-08-11

Thank you for these nice explanations! I fully agree that studying and recovering dynamics from biased simulations is an important problem and appreciate any ML work done for it. I think the paper should be accepted. I appreciate the chignolin experiment. 1. Minor: Could you put your work a bit into context with "Implicit Transfer Operator Learning: Multiple Time-Resolution Surrogates for Molecular Dynamics" https://arxiv.org/abs/2305.18046 . They only learn the transfer operator. However, they also provide experiments on Chignolin. Are there inherent reasons why learning the generator might scale worse than learning the operator? My main concern (maybe it was naive) was in thinking that CVs (which you are learning by learning the generator and using its eigenfunctions) main or only value is in using them for further biased simulations. Thus, the question of why one biased simulation with your CVs would be better than the first biased simulation. This also seems valuable, and you suggest it in the general rebuttal, but do not try it. \ Important:\ However, as far as I understand, you suggest that I was wrong in assuming that learning the generator is only valuable for obtaining CVs and using them for biasing simulations. Learning them is also valuable to e.g. recover the free energy surface from a biased simulation or to recover dynamical quantities about transitions - please let me know if this understanding is correct. (sorry for the delayed reply - I will be quicker to respond in the remaining days)

Authorsrebuttal2024-08-12

Acknowledgement of reviewer's comments

We thank the reviewer for appreciating our rebuttal and suggesting our paper for acceptance. In the following we address the reviewer's additional questions. - __Concerning the reference.__ Thanks for bringing this paper to our attention, we will include it in the revision. In this work the authors build their model from D.E. Shaw dataset formed by a long trajectory containing already all the necessary information for training, without the need for biasing. They learn a transition kernel (i.e. a conditional probability to go from $X_t$ to $X_{t+\Delta t}$), which, as discussed in Sec. 3, has an inherent difficulty to be adapted to biased simulations. - __Scaling of generator algorithms.__ First, note that transfer operator (TO) based methods apply only to equally spaced data and that the sampling frequency $1/\Delta t$ must be high enough to distinguish all relevant time-scales in the dynamics. Otherwise, since TO eigenvalues are $e^{\lambda_i \Delta t}$, small spectral gaps complicate learning (see Thm. 3 [23]). Conversely, our IG method, which uses gradient information, is time-scale independent, handles irregularly spaced measurements, and does not rely on time discretizations. This important aspect allows one to learn from biased simulations without notorious time-lag bottlenecks inherent to TO methods. Alas, as there is no free lunch, this incurs higher computational complexity. However, his complexity is to an extent mitigated through our representation learning, by exploiting automatic differentiation tools in deep learning. - __IG’s eigenfunctions.__ Indeed, the learned eigenfunctions of the IG can be used as CVs, which in fact are optimal CVs, as discussed in the general reply, see [A]. Moreover, the leading eigenfunctions of IG encode the true dynamics. Hence, once they are properly learned, no more biasing is needed. For example, they can be used to infer the transition mechanism, see for instance figure 2, c) in the alanine dipeptide experiment, where we show that even with few transitions, we manage to recover a linear relationship between $\theta$ and $\phi$. Another important aspect is that one can build good approximations of all transfer operators $\cal{T}_{\Delta t}$, $\Delta t>0$, from the leading eigenpairs of IG, enabling forecasting of system’s observables see e.g. [21]. If there are any remaining concerns and/or questions, we are happy to discuss more.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC