Learning the Infinitesimal Generator of Stochastic Diffusion Processes

We address data-driven learning of the infinitesimal generator of stochastic diffusion processes, essential for understanding numerical simulations of natural and physical systems. The unbounded nature of the generator poses significant challenges, rendering conventional analysis techniques for Hilbert-Schmidt operators ineffective. To overcome this, we introduce a novel framework based on the energy functional for these stochastic processes. Our approach integrates physical priors through an energy-based risk metric in both full and partial knowledge settings. We evaluate the statistical performance of a reduced-rank estimator in reproducing kernel Hilbert spaces (RKHS) in the partial knowledge setting. Notably, our approach provides learning bounds independent of the state space dimension and ensures non-spurious spectral estimation. Additionally, we elucidate how the distortion between the intrinsic energy-induced metric of the stochastic diffusion and the RKHS metric used for generator estimation impacts the spectral learning bounds.

Paper

References (57)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer D2P38/10 · confidence 4/52024-07-11

Summary

This paper proposes a relevant and sound approach for learning self-adjoint SDE generators via operator learning techniques. The paper includes a compactification, a novel prior knowledge inclusion, and first-of-its-kind statistical learning guarantees that extend the known ones from discrete Markov processes.

Strengths

1. The paper is easy to read, with suitable mediation throughout. 2. Inserting known diffusion effects into the Dirichlet forms and using the resolvent, provides an elegant way of including prior knowledge and well-posed estimation. 3. These are the most complete statistical learning guarantees for spectra of self-adjoint generators (albeit reliant on partial knowledge).

Weaknesses

1. Unclear extensions to entirely data-driven regimes (no partial knowledge) and requiring sampling from an invariant distribution. 2. Experiments only consider toy examples that validate theory but do not showcase the practicality of the approach for more general cases. 3. I am not sure that contribution **4)** is well supported by the current presentation, primarily due to hard-to-parse figures. A revised comparison and legends with sharper plots could make the results much easier to understand.

Questions

1. *Are Dirichlet forms commonly deduced from the diffusion part of the SDE?* It would help assess the ease of knowing these forms beforehand for readers unfamiliar with the notion. 2. *Is it reasonable to think you could deal with a bias of the Dirichlet form using imperfect diffusion knowledge?* 3. *Could you highlight what practical gains/losses are introduced w.r.t. discrete-time setting?* I would guess you need less data due to diffusion priors and not requiring as small of a discretization. On the other hand, IG regression complexity seems higher for high-dimensional systems. 4. *Can a sample complexity gain be recognized using prior knowledge in this form for your theoretical results?* Could the performance gain w.r.t. TO learning be recognized using the derived bounds? 5. *How arbitrary is the choice of $\mu$?* 6. *How does the time-sampled data come into play, and does the sampling rate of the data influence any aspect of the approach?* Does sampling the invariant distribution + Dirichlet form allow you to avoid time-derivative or "1-step" dynamics observations (e.g., compared to [A])? Based on the current writing, it is not immediately clear to me (but I might have missed something). 7. *How do you compare (using full knowledge) to classical numerical methods using known drift and diffusion? Would your sample efficiency be better than classical FEM for a given precision?* You mention how, for known operators (via drift and diffusion knowledge), there would be no spuriousness in estimating eigenpairs. This is perhaps of separate interest. [A] - Meng, Y., Zhou, R., Ornik, M., & Liu, J. (2024). Koopman-Based Learning of Infinitesimal Generators without Operator Logarithm. http://arxiv.org/abs/2403.15688

Rating

8

Confidence

4

Soundness

4

Presentation

4

Contribution

3

Limitations

1. *Possibly strong prior knowledge requirements.* 2. *Limited to self-adjoint generators.* 3. *Experiments are on 1D toy examples.*

Reviewer D2P32024-08-11

Rebuttal Reply for Submission18874 by Reviewer D2P3

I thank the authors for the extensive rebuttal in the general reply and comprehensively addressing my review comments. Also, it is great to see the authors put in the effort to deliver adjusted theory as well as run new experiments. Many important aspects of the paper got clearer: limits of FEM and TO approaches, fully-data driven regime, energy functional rationale and shift parameter. Accordingly, I have increased my overall score. I expect this rebuttal content to be incorporated as part of a final submission.

Authorsrebuttal2024-08-12

Acknowledgement to the reviewer

We would like once more to sincerely thank the reviewer for their deep questions, which inspired us to make these additional steps and improve our work. We are happy that our rebuttal is appreciated and we commit to incorporate it in the revised manuscript.

Reviewer PUzn7/10 · confidence 3/52024-07-13

Summary

The paper considers a time-homogeneous Stochastic Differential Equation (SDE) with known diffusion part and known or unknown drift. The problem is to find properties of this equation from known data, in particular, to find a (low-rank) representation of its infinitesimal generator (IG). For this purpose, the resolvent of this operator is considered and the reduced-rank estimator in reproducing kernel Hilbert spaces (RKHS) together with the energy-based risk metric are used. Accuracy estimates for the found approximation to IG and time-complexity are given. Several model numerical experiments are conducted for proof-of-concept.

Strengths

- Detailed appendix with necessary reference material - Development of SDE theory with exact estimates on the result obtained

Weaknesses

- The paper uses quite complex concepts and may be difficult to understand in a 9 page format. - There are not enough practical examples of the application of these results to machine learning Minor The x-axis in Figure 1 (a) is not signed L120 comma is not needed L124 space after dot is missing L206: colon missing at the end of the line L215 A new line is redundant (`this work we focus on the case when`) L275 space after dot is missing L915 "anan" -> "and"

Questions

- Can your method be extended to the case when neither drift nor diffusion coefficients are known?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Limitation are described in the paper. The paper is mostly theoretical.

Reviewer 6wyC6/10 · confidence 2/52024-07-15

Summary

In this paper, the authors consider the problem of learning the infinitesimal generator a Stochastic Diffusion Process (SDP). Compared to existing approaches such as [1] they tackle the unbounded nature of the generator by introducing a novel statistical framework which is based on the Dirichlet form associated with the SDP. In this framework, the authors estimate the resolvent of the generator which can be approximated with finite-rank operators. They consider a regularized and rank-truncated version of the regression loss. Theorem 1 provides a way to compute the eigenvalues and eigenvectors of the estimated operator. Spectral learning bounds are derived in Theorem 2. The theoretical results of the paper are completed with three examples: learning the overdamped Langevin generator for a one-dimensional four well potential, the overdamped Langevin generator for a Muller Brown potential and finally for a Cox-Ingersoll-Ross process (CIR process). [1] Cabannes et al. (2023) -- The Galerkin method beats Graph-Based Approaches for Spectral Algorithms

Strengths

* Even though the paper is mathematically heavy and contains a lot of notation, I think it is well-written. I appreciate that (almost) all the assumptions needed before stating Theorem 2 are clearly laid out and explained. I also appreciate the rigour shown in the paper. * The obtained results regarding the spectral bounds are interesting and provide a strong theoretical grounding for the method. * From a methodological point of view the time complexity can be reduced compared to [1,2,3] if the rank is low. I think this is one of the strength of the method, although I would have appreciated more details on the choice of the hyperparameters and their interactions (see below). [1] Cabannes et al. (2023) -- The Galerkin method beats Graph-Based Approaches for Spectral Algorithms [2] Hou et al. (2023) -- Sparse learning of dynamical systems in RKHS: An operator-theoretic approach [3] Pillaud-Vivien et al. (2023) -- Kernelized Diffusion Maps

Weaknesses

Before expressing my main concerns regarding this paper, I want to emphasize that I'm not an expert in this domain and therefore my understanding of the main competitors and methods used in the paper is lacking. * My main concern is regarding the experiments. I know this is a theoretical paper but a new statistical framework and a new learning procedure should be clearly validated against competing methods like [1,2,3]. Even though some qualitative conclusions are drawn I would have liked to see a more extensive study (for different choices of hyperparameters for both the target Langevin diffusion and the method introduced by the authors). * Indeed, there are little details in the paper about how to choose the crucial parameters of the algorithm like $\mu$, $r$ and $\gamma$. Establishing the robustness of the procedure regarding these hyperparameters seem crucial to validate the methodology. * As mentioned before, even though I appreciate the rigour of the paper it is notation heavy. Having to constantly rely on the (massive) table of page 13 to remember the notation hindered my understanding of the paper. Is there a way to either reduce the notational load or to incorporate a lightweight version of Table 2 in the main paper? This would greatly help the reader. * All assumptions are commented except (KE). What are the main limitations of this assumption? I am not familiar with the domain so a light explanation in the main text would also be appreciated. * Is the assumption of l.188 that there exists a Dirichlet operator realistic? It is true for Langevin operator and CIR (which already covers quite a lot of ground) but apart from these models? Does this impose anything on the diffusion? Some minor remarks: * l.178 "Importantly, this energy functional can be empirically estimated from data sampled from π, whenever full knowledge, that is drift and diffusion coefficients of the SDE (1), or partial knowledge," -- This sentence is not really clear to me. How can one leverage the drift and diffusion coefficient here? * l.215 typo (spurious new line) [1] Cabannes et al. (2023) -- The Galerkin method beats Graph-Based Approaches for Spectral Algorithms [2] Hou et al. (2023) -- Sparse learning of dynamical systems in RKHS: An operator-theoretic approach [3] Pillaud-Vivien et al. (2023) -- Kernelized Diffusion Maps

Questions

See weaknesses

Rating

6

Confidence

2

Soundness

3

Presentation

3

Contribution

2

Limitations

Limitations are addressed in Section 7 ("Conclusion")

Reviewer nFzp5/10 · confidence 3/52024-07-26

Summary

This paper discusses learning the generator of stochastic diffusion processes in reproducing kernel Hilbert space. In particular, section 2: background on the generator, Dirichlet form, energy, learning in RKHS, empirical risk in the Hilbert-Schmidt norm section 3: introduce an energy-based risk functional for resolvent estimation which leads to spectral estimation section 4: empirical risk minimization with Tikhonov regularization and rank constraints.

Strengths

The paper is well-written and has a thorough discussion on the background and the comparison of previous work. It proposes a novel energy risk functional and a reduced-rank estimator with dimension-free learning bounds. It also establishes spectral learning bounds for generator learning.

Weaknesses

The weaknesses are mainly in the motivation (please see question 1 below) and the experiment (see questions 3 and 4 below). The authors did a great job in the theoretical comparison with previous work on TO and IG learning, however, the empirical comparison focuses on the spuriousness of eigenvalues and the recovery capability of the metastable states. To sufficiently demonstrate the practical advantage of the proposed method, other comparisons of the learned dynamics, phase transition, and invariant measure should also be included.

Questions

1. In the paper, 'full knowledge' refers to given drift and diffusion coefficients, and 'partial knowledge' refers to the given diffusion coefficient and energy/Dirichlet operator. In practice, is it necessarily easier to obtain diffusion coefficient and energy/Dirichlet operator than to obtain drift and diffusion coefficients? In a fully data-driven scenario, one can first estimate drift and diffusion coefficients, and then use physics-informed method with 'full knowledge' obtained. Does this make the 'full knowledge' methods cited in the paper more flexible than the 'partial knowledge' method proposed? 2. As transfer operator A_t and the infinitesimal generator L satisfy the relation A_t = e^{Lt}, is there a consistency between the learned TO from cited TO methods with the learned IG from the proposed method? 3. What are the kernels used for the 3 examples in experiments? How sensitive is the experimental result to the kernel selection, regularization, and rank parameter? 4. Can the authors provide the learned dynamics of the 3 examples compared to the true dynamics, the invariant measure, and time scales?

Rating

5

Confidence

3

Soundness

3

Presentation

4

Contribution

3

Limitations

Limitations have been discussed.

Reviewer PUzn2024-08-12

I thank the authors for their explanations. After reading the global response and the discussion with other reviewers, I believe that in general my concerns were addressed. I raise my score to 7 "Accept".

Authorsrebuttal2024-08-12

Acknowledgement to the reviewer

We would like once more to thank the reviewer for their comments, which inspired us to make additional steps and improve our work.We are happy that our rebuttal was helpful and we commit to incorporate it in the revised manuscript.

Authorsrebuttal2024-08-12

Acknowledgement to the reviewer

We would like once more to thank the reviewer for their comments, which inspired us to make additional steps and improve our work.We are happy that our rebuttal was helpful and we commit to incorporate it in the revised manuscript.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC