Summary
This paper utilizes theory developed around random projections to develop fast/efficient numerical methods for approximating the Koopman mode decomposition of data sets with long-time trajectories, that additionally come with learning rates that are optimal.
Strengths
1. This paper developed new fast/efficient algorithms for computing the Koopman spectra, the effectiveness of which was demonstrated on very large data sets with long-time trajectories. These include a traditional benchmark (Lorenz '63), as well as molecular data sets, which are a growing area of interest to the machine learning and Koopman communities. Additionally, the molecular data sets were of a size that could not have been studied previously.
2. This paper provided learning rates for their new algorithms, which they show to be optimal.
3. This paper discusses how different methods for computing the Koopman mode decomposition (e.g., Nystrom RRR vs. PCR) differ in their learning rate dependencies.
4. The paper was well-written.
Weaknesses
1. Motivation for why very long-time trajectories of data would be needed for some systems was missing from the Introduction. This was partially discussed at the end in Sec. 5 "Molecular dynamics datasets", but developing this more is important. Additionally, the title and some places in the text describe the sketching approach as being useful for "large scale dynamical systems". As I understand it, the methods are beneficial for "long-time trajectory dynamical systems". Of course, the molecular examples shown have both large dimension and long-time trajectories, but the title and text should be corrected to accurately emphasize that it is long-time trajectory dynamical systems that this method is useful for.
2. The discussion on how Nystrom PCR and RRR differ in their learning rates was very interesting. Including more on this (maybe in table form) would be nice. Additionally, it would be helpful to see all 3 methods compared, so KRR should be included in Figure 1. Additionally, the reason for using just RRR in Figs. 3 and 4 should be noted.
3. I found it a little confusing that going from Eq. 8 to Eq. 9, the P_Y \hat{C}_{YX} P_X became K^\dagger_{\tilde{Y}, \tilde{Y}} K_{\tilde{Y}, Y} K_{X, \tilde{X}}, but that going from Eq 10 to Eq 11 the P_Y \hat{C}_{YX} also led to K^\dagger_{\tilde{Y}, \tilde{Y}} K_{\tilde{Y}, Y} K_{X, \tilde{X}}. I assume this has to do with the way the \dagger was absorbed into the [ \cdot ]_r?
4. How the random $\tilde{\Phi}$ are generated should be explained in the text.
Minor comments:
1. Mezic, 2005 should be cited when discussing the Koopman mode decomposition and Budisic et al. 2012 should be cited when discussing how Koopman mode decomposition has been leveraged in the past (line 95).
2. Abbreviations should be explained before using them (DMD, tICA, VAC, etc.).
3. The box plot denoting the sKAF data in Fig. 2 has a strange shape in its lower box (an inverted U). Why is this?
4. It would be helpful and appropriate to have citations when noting that "An important application of Koopman operator theory is in the analysis of molecular dynamics".
5. The colors of Fig. 3 should be explained in more detail.
6. It might be helpful to have a line across the 3 states in Fig. 4 to show how Eigenfunction 1 has linear separation between State 0 and the others, and Eigenfunction 2 has linear separation between State 1 and the others.
7. There are a few typos:
i. "allows to deploy" (line 1)
ii. "as theirs exact" (line 68)
iii. "to study large class" (line 90)
iv. "]0, 1]" (Assumption 4.3)
v. "]0, \tau]" (Assumption 4.4)
Questions
To summarize, the Weakness section above,
1. Why would one want to use very long-time trajectory data when computing the Koopman mode decomposition? (You explain this some, but making it more explicit would be helpful)
2. What are the differences between KRR, PCR, and RRR in their performance on the different data sets studied (at least the Lorenz '63)?
3. How are the random $\tilde{\Phi}$ are generated?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The authors do a sufficient job explaining their limitations (the assumptions they make) and what kinds of data sets their proposed methods work best on. The only exception to this is that it should be clarified that the proposed methods improve Koopman mode decomposition estimates of long-time trajectory data and not large dimension data.