Summary
This paper discusses learning the generator of stochastic diffusion processes in reproducing kernel Hilbert space. In particular,
section 2: background on the generator, Dirichlet form, energy, learning in RKHS, empirical risk in the Hilbert-Schmidt norm
section 3: introduce an energy-based risk functional for resolvent estimation which leads to spectral estimation
section 4: empirical risk minimization with Tikhonov regularization and rank constraints.
Strengths
The paper is well-written and has a thorough discussion on the background and the comparison of previous work. It proposes a novel energy risk functional and a reduced-rank estimator with dimension-free learning bounds. It also establishes spectral learning bounds for generator learning.
Weaknesses
The weaknesses are mainly in the motivation (please see question 1 below) and the experiment (see questions 3 and 4 below). The authors did a great job in the theoretical comparison with previous work on TO and IG learning, however, the empirical comparison focuses on the spuriousness of eigenvalues and the recovery capability of the metastable states. To sufficiently demonstrate the practical advantage of the proposed method, other comparisons of the learned dynamics, phase transition, and invariant measure should also be included.
Questions
1. In the paper, 'full knowledge' refers to given drift and diffusion coefficients, and 'partial knowledge' refers to the given diffusion coefficient and energy/Dirichlet operator. In practice, is it necessarily easier to obtain diffusion coefficient and energy/Dirichlet operator than to obtain drift and diffusion coefficients?
In a fully data-driven scenario, one can first estimate drift and diffusion coefficients, and then use physics-informed method with 'full knowledge' obtained. Does this make the 'full knowledge' methods cited in the paper more flexible than the 'partial knowledge' method proposed?
2. As transfer operator A_t and the infinitesimal generator L satisfy the relation A_t = e^{Lt}, is there a consistency between the learned TO from cited TO methods with the learned IG from the proposed method?
3. What are the kernels used for the 3 examples in experiments? How sensitive is the experimental result to the kernel selection, regularization, and rank parameter?
4. Can the authors provide the learned dynamics of the 3 examples compared to the true dynamics, the invariant measure, and time scales?
Limitations
Limitations have been discussed.