Parametric model reduction of mean-field and stochastic systems via higher-order action matching

The aim of this work is to learn models of population dynamics of physical systems that feature stochastic and mean-field effects and that depend on physics parameters. The learned models can act as surrogates of classical numerical models to efficiently predict the system behavior over the physics parameters. Building on the Benamou-Brenier formula from optimal transport and action matching, we use a variational problem to infer parameter- and time-dependent gradient fields that represent approximations of the population dynamics. The inferred gradient fields can then be used to rapidly generate sample trajectories that mimic the dynamics of the physical system on a population level over varying physics parameters. We show that combining Monte Carlo sampling with higher-order quadrature rules is critical for accurately estimating the training objective from sample data and for stabilizing the training process. We demonstrate on Vlasov-Poisson instabilities as well as on high-dimensional particle and chaotic systems that our approach accurately predicts population dynamics over a wide range of parameters and outperforms state-of-the-art diffusion-based and flow-based modeling that simply condition on time and physics parameters.

Paper

Similar papers

Peer review

Reviewer rAtu5/10 · confidence 4/52024-07-06

Summary

The authors present a framework for learning a time-dependent gradient vector field (parametric model) that describes how the particle evolves in the state space. The idea is based on the action-matching framework (https://arxiv.org/abs/2210.06662), which learns a time-dependent gradient field that interpolates between marginals at different times in a simulation-free manner. The authors propose using high-order quadrature rules for evaluating nested integrals in the objective. They demonstrate the forecasting ability of their framework on diverse examples in high dimensions and show that it outperforms other diffusion-based and flow-based models in inference runtime and accuracy

Strengths

Developing efficient and high-fidelity surrogate models is of great importance to the computational science and engineering community. The authors take a step forward in this direction by learning the effective population dynamics of the underlying physical process. The paper is well-written, and the related works are adequately referenced.

Weaknesses

The objective used by the authors is based on the action-matching paper, now adapted to include the dependence on the parameters (\mu). From a methodological standpoint, the authors' only important contribution is using a higher-order quadrature scheme to discretize the nested integrals

Questions

General comments: i) Maybe, I am not understanding things here.. Can the authors elaborate a bit more on the O(K\tau) complexity? If for instance, I train a velocity field using continuous normalizing Flow (simulation-based) using empirical marginals at different times. I can then integrate the learned velocity in time to get samples, in the same way, the authors draw samples from the gradient field through Langevin dynamics. Where is the extra \tau in complexity coming from? ii) Or, is this framework a simulation-free strategy to train a Continuous Normalizing Flow but adapted for multiple snapshots? This was not so clear in the paper. iii) The authors learn a gradient field as a function of parameters and time. Can I authors comment a bit on the stability of the learned velocity during inference time? Specific comments i) In Figure 5, a color bar showing the scales would help with the interpretation. ii) In Figure 4, a label to show what the two curves (blue and orange) are is needed. What is the ground truth and what is the prediction?

Rating

5

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

-

Reviewer gej19/10 · confidence 4/52024-07-09

Summary

The authors focus on learning models for population dynamics of parameterized physical systems that exhibit stochastic and mean-field effects in time. To do so, the authors leverage the Benamou-Brenier formula to learn gradient fields that transport the probability density as time evolves, and that enable the generation of sample trajectories reflecting the dynamics of the population. Numerical experiments show compeling results and state-of-the-art performance in high-dimensional particle systems and in chaotic systems.

Strengths

The contributions of the paper are novel, and the presentation is very clear. I commend the author's attempts at making the paper reasonably self-contained by adding context in Appendices C through E. Each of the theoretical developments are clearly motivated. Finding a vector field as proposed enables the interpolation of probability measures at different times. This notion has broad impacts beyond those stated in the paper. It is also interesting that the authors illustrate the importance of using a proper quadrature in time as opposed to random sampling. The experiments are also compelling, in particular the ability to resolve the low probability connection between two high probability regions as shown in Fig. 5.

Weaknesses

To this reviewer the paper does not have any evident weakness. However, when introducing a high-order quadrature in time, which improves the performance, the authors do not elaborate on why one would, or would not, expect further improvements if similar quadrature rules were used in the spatial variable or in the physical parameters.

Questions

I have the following minor questions: - In (6) it seems that the integrand should be evaluated on $s_t$ instead of $s$ for consistency with (4). - In (7) there seems to be a parenthesis missing that factors out the density $\rho_{t,\mu}$.

Rating

9

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

The authors adequately address the limitations of their work in Section 4. These are mostly technical. For instance, they require a high sampling rate in time, which is not always available in practice. This limits the applicability of the method, but does not detract from their main contribution.

Reviewer zRLH4/10 · confidence 3/52024-07-13

Summary

This paper develops models of population dynamics in physical systems that exhibit stochastic and mean-field effects, influenced by physics parameters. The goal is to create models that can efficiently predict system behavior as alternatives to classical numerical methods. By utilizing the Benamou-Brenier formula from optimal transport and action matching, the approach involves solving a variational problem to infer gradient fields that approximate population dynamics. These gradient fields enable the generation of sample trajectories that mimic physical system dynamics under various physics parameters. The study highlights the importance of combining Monte Carlo sampling with higher-order quadrature rules for accurate estimation and stable training. The models are demonstrated to perform well on Vlasov-Poisson instabilities and high-dimensional particle and chaotic systems, outperforming state-of-the-art diffusion-based and flow-based models that rely solely on time and physics parameters.

Strengths

1) Efficient parametric model reduction: The model is demonstrated to reduce inference runtime significantly compared to standard diffusion- and flow-based models by leveraging minimal-energy vector fields. 2) Accurate and stable dynamics learning: It captures the coupling over time steps accurately, using higher-order quadrature schemes for estimating time integrals, which enhances training stability. 3) High accuracy and reduced runtime: Achieves error rates comparable to state-of-the-art methods while reducing inference runtime by 1-2 orders of magnitude.

Weaknesses

1) The model assumes access to a dense set of time points for the Gauss-Legendre quadrature, which may not be applicable when only a few time samples are available. This was already mentioned as a limitation of the current work. As this forms the main part of the approach, the practical benefit would be limited. 2) Regarding the vector field complexity, the model seeks a vector field that minimizes kinetic energy, but in some cases, this may be more complicated than other vector fields that produce the same population dynamics. Examples include situations where the minimal-energy field varies with time, making it challenging to determine the appropriate energies to use for different problems.

Questions

1) Authors aims to learn population dynamics $\rho_{t,\mu}$ instead of learning the dynamics of individual trajectories $t\to X_{t,\mu}^i$ for all $i$/. What are the assumptions on model so that it admits density? If not, would this work apply to the situation under weak convergence where no density is assumed. Moreover, it is not clear why the equations (5) to (7) should admit solutions in the strong forms. 2) Parametrizing the $s_{t,\mu}$ with weight modulation as in CoLoRA [38] only applies to deterministic time-dependent dynamical systems. However, one goal of the work was stated as learning the population dynamics should allow for seamless treatment on deterministic and stochastic systems. How the latter can be handled with the corresponding parametrization is unclear. The forms of the layers assumes some low-rank structure, and only the weight modulations $\phi$ is assumed to depend on time and parameters. As such, are stochastic systems omitted from the problem definition unlike the initial motivation? 3) Solving eq (7) from data is challenging and prone to potential numerical issues. As a remedy, a combination of higher order numerical quadrature and MC sampling strategy is proposed. However, the details of this approach would be better to provide with a stability analysis on the training as mentioned in the text.

Rating

4

Confidence

3

Soundness

2

Presentation

2

Contribution

2

Limitations

Yes, limitations of their work is adequately addressed. No potential negative societal impact of their work is identified.

Authorsrebuttal2024-08-13

> I believe all of these will require additional rewrite in multiple places. We thank the reviewer for raising these concerns. We are confident that all of the concerns are addressed in the paper, which we concisely summarize in this response. The reviewer’s comments will be helpful for guiding the final revision of this paper (if accepted) to make our points even clearer. > it is yet unclear to me how the proposed model improves the results and helps the stabilization of the training as claimed. We summarize here the reasoning for why "the proposed model improves the results and helps the stabilization of the training", which is also provided in the paper but we will make this clearer in a revision, if the paper gets accepted: 1. **Gauss-Legendre quadrature results in a more accurate estimate of the time integral when compared to Monte Carlo.** This is because it is a higher-order numerical scheme (more precisely, it allows integrating higher-degree polynomials exactly) that leads to lower quadrature errors than Monte Carlo. Besides this theoretic argument, the difference in the quadrature error can be seen empirically in our experiments in, e.g., Figure 2. (See Section 2.3.) 2. **A more accurate estimate of the time integral results in a more stable and accurate estimate of the loss.** This is supported numerically by Figure 2 (left) which shows extremely high variability in the estimates of the loss for AM and low variance estimates of loss for our proposed HOAM. This also agrees with standard results from statistical learning theory where the deviation of the empirical risk (the estimate of the loss function) from the true risk (the true loss function) directly enters in bounds of generalization errors. 3. **A more accurate estimate of the loss makes solving the optimization problem tractable.** This is supported by the numerical results in the paper, explicitly in Figure 2 (middle) and shown to be essential in global response, All this is also described in detail in the main text of our paper in Section 2.3 and numerically supported by Figure 2 and the global response. For example see line 200: "Plain uniform sampling over the data set for mini-batching can lead to poor estimates of the loss." > If I understand correctly, the HOAM combines higher order numerical quadrature with MC sampling to solve (7) in the direction of time by employing a Gauss-Legendre quadrature. How can this be achieved numerically? 1. As referenced in the main text of the paper the empirical loss is given in the Appendix in section A. This is written down as a discrete summation which can be easily implemented numerically. 2. For the Gauss-Legendre quadrature, the key algorithmic step is determining the roots of the Legendre polynomials, which is implemented (for example) in scipy (scipy.special.roots_legendre()). We also provide the code of our implementation, but the link is redacted for now per submission policy. > It would be also useful to motivate the choice of the dynamical systems in the numerical experiments? Why the examples are relevant in this context is unclear. 1. From a surrogate modeling perspective, the systems that we consider are exceedingly challenging because they are high dimensional, chaotic and/or stochastic. Surrogate modeling for such systems is in its infancy because the point-wise approximations of traditional surrogate modeling techniques are meaningless. We reference these attempts in detail in the introduction and literature review. 2. We make careful efforts to choose problem setups which are experimentally meaningful to practitioners (see for example Appendix B.2 that describes that the data are obtained from code that are used by plasma physicists). The physical relevance of our problem setups are detailed in the numerous sources we cite, see [7, 21, 34, 42, 87 Sec 2(b)(i), 58]. > A clear framing of the problem setup and all the assumptions on the model would help the reader to better understand the results. 1. The problem setup is described in the Introduction section on page 1: “Given a data set of samples [...] we aim to learn a dynamical-system reduced model to rapidly predict samples that approximately follow the same law [...]” 2. When we make mathematical statements, we provide assumptions. For example, for the statement of ‘uniqueness’ on page 4 we assume that the density is positive and that the source term integrates to zero. 3. We provide an extensive appendix C-E on the connection to optimal transport and additional literature. We hope that we have been able to demonstrate that most of these concerns are addressed in the main text of the paper. If the reviewer does not have any additional concerns, we would greatly appreciate it if they would consider revising their score.

Reviewer rAtu2024-08-12

I thank the authors for their detailed responses and clarification. I will retain my rating.

Authorsrebuttal2024-08-13

Thank you for the comments. If there is any other information we can provide, please let us know.

Reviewer zRLH2024-08-13

Thank you for the rebuttal and the additional results. After reading your responses, I still have same concerns regarding the contribution of the paper. My main issue is that while the authors claim that higher-order action matching (HOAM) is better than the baseline action matching (AM) on some parametric dynamical systems, it is yet unclear to me how the proposed model improves the results and help the stabilization of the training as claimed. A clear framing of the problem setup and all the assumptions on the model would help the reader to better understand the results. I also find the implementation a bit unclear. If I understand correctly, the HOAM combines higher order numerical quadrature with MC sampling to solve (7) in the direction of time by employing a Gauss-Legendre quadrature. How can this be achieved numerically? It would be also useful to motivate the choice of the dynamical systems in the numerical experiments? Why the examples are relevant in this context is unclear. Hence, I believe all of these will require additional rewrite in multiple places. Therefore, I will keep my initial rating for now.

Reviewer gej12024-08-13

I thank the authors for their responses. I have the following additional question. In the pdf file you attached, Fig. 3 shows the relative error in mean and the caption states that "we see that HOAM remains relative stable, while AM increases in error." However, HOAM is the _blue_ line which not only is above the orange line (representing AM) but also seems to increase faster as $T$ increases. Is this a typo?

Authorsrebuttal2024-08-13

Yes this is a typo. We thank the reviewer for catching this. We apologize for the confusion and will update the pdf accordingly. In all cases HOAM does outperform AM.

Reviewer gej12024-08-13

Thanks for the clarification.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC