Learning diffusion at lightspeed

Diffusion regulates numerous natural processes and the dynamics of many successful generative models. Existing models to learn the diffusion terms from observational data rely on complex bilevel optimization problems and model only the drift of the system. We propose a new simple model, JKOnet*, which bypasses the complexity of existing architectures while presenting significantly enhanced representational capabilities: JKOnet* recovers the potential, interaction, and internal energy components of the underlying diffusion process. JKOnet* minimizes a simple quadratic loss and outperforms other baselines in terms of sample efficiency, computational complexity, and accuracy. Additionally, JKOnet* provides a closed-form optimal solution for linearly parametrized functionals, and, when applied to predict the evolution of cellular processes from real-world data, it achieves state-of-the-art accuracy at a fraction of the computational cost of all existing methods. Our methodology is based on the interpretation of diffusion processes as energy-minimizing trajectories in the probability space via the so-called JKO scheme, which we study via its first-order optimality conditions.

Paper

Similar papers

Peer review

Reviewer vFxe7/10 · confidence 3/52024-06-30

Summary

This paper considers learning diffusion dynamics from observational data of populations over time, identified as learning the energy functional in Equation 3. Past research has confronted this inverse problem via complex bilevel optimization, limited to potential energies. This paper proposes an alternative model JKOnet* that can work with potential, internal, and interaction energies, efficiently minimizes a quadratic loss instead of a complex bilevel optimization, has much lower computational complexity, and out-performs baselines in simulations. A variant for linearly parameterized functionals has a closed form solution. The paper's new method reconsiders the JKO scheme using first-order optimality conditions, resulting in decompose the problem into first computing optimal transport plans between adjacent populations and then optimizing a loss for fixed plans.

Strengths

- Inferring diffusion dynamics from observational data is a difficult and significant problem for which this paper appears to provide a solid contribution. The paper substantially improves upon JKOnet in terms of multiple directions: better performance (Figure 3), simpler optimization objective (Equation 11), better scalability and efficiency (e.g. Table 1, Section 4.2), and improved generality (Table 1, Section 4.3). These dimensions are analyzed in experiments across a range of different energy functionals, where the gains are shown in log-scale displaying orders of magnitude improvement. The paper makes a convincing argument for using JKOnet* over JKOnet. - The methodology appears quite strong, well-motivated, and original, with solid intuition given by the authors throughout the paper.

Weaknesses

Minor weaknesses: - While the results are strong, occasionally the language feels too imprecise. For example, "runs at lightspeed" seems inaccurate compared to "runs very efficiently". The authors also mention that they rely upon weeks-old advancements in optimization in the abstract which seems unneeded. - The paper is generally very well-written except for the introduction which could use editing. It introduces a lot of terminology and details from past research. Similarly, Figure 1 is referenced multiple times including in the introduction but it was hard to understand until after reading Section 3. - The construction of the optimal transport plans does not seem to be included in the computational complexity comparisons. While this is computed once for JKOnet*, it is additional expense over JKOnet.

Questions

1. What is JKOnet_l in Table 1? 2. In Section 4.2, the authors conclude that JKOnet* is well-suited for high-dimensional tasks. Does this include computing the optimal transport maps? 3. The discussion in Figure 3 in the text focuses primarily on the speed improvement, yet the performance gains are also quite large, including seemingly between JKOnet* and JKOnet*_l. Can the authors comment on why the linear parameterization was useful in their experiments?

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

4

Limitations

Limitations are adequately addressed in Section 5

Reviewer fc3q8/10 · confidence 4/52024-07-10

Summary

The authors study diffusion processes from the perspective of Wasserstein gradient flows. Based on the recent fixed-point characterisation for Wasserstein proximal operator methods, they introduce Jordan-Kinderlehrer-Otto (JKO) type methods for learning potential and interaction energies that govern the diffusion process. Such methods are assuming that a sample of the population distribution at each time step is at hand (not necessearily obtained by tracking individual particles) implying important applications across various fields. While theoretical novelties are present (w.r.t. paper [26] that lies in the foundation of this work), the main contribution is the overall methodology for learning diffusion processes.

Strengths

Paper is, besides minor issues reported bellow, excellently written - very clear, precise and intuitive with well balances technical details between main text and the appendix. Existing ideas are neatly combined to obtain significant improvements of the JKO-type methods and extensive empirical evaluation is presented. The proofs seem correct and well-written.

Weaknesses

While I do not find important weaknesses, I feel that next several small issues can be addressed to further improve readability: 1. When addressing content presented in the appendix it would be good to refer to the section, e.g. see Figure 6 in Appendix A. 2. It would be good to say what $\rho_t$ is in Example 2.1 3. While Table 3.1 reports per-epoch complexity for all the methods, it would be important to note that JKOnet$^*$ have additional computational complexity for solving $T$ OT problems of size $N$ in $d$-dimensions. Detailed remark on the initial computational complexity, depending of the algorithm used, should be reported. 4. In Section 4 it would be helpful to introduce the problems, that is to better explain the task of each experiment and the role of functionals ($V(x)$ ?!) appearing in Appendix F. Maybe giving an example on Styblinski-Tang functional appearing in Figures 2, 3 and 4, and then referring to other ones by their names and/or reference equations.

Questions

1. In the implementation of the method, a priori computed optimal transport plans are obtained by solving entropy-regularised OT via Sinkhorn-type algorithms or some other methods? 2. What do you think about the applications and/or limitations of the JKOnet$^*$ for the setting of long-trajectories to infer the behaviour in equilibrium, e.g. detection of meta-stable states of Langevin dynamics?

Rating

8

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

Limitations are addressed adequately.

Reviewer vNRb6/10 · confidence 3/52024-07-10

Summary

This paper introduces JKOnet*, a new method for learning diffusion processes from data. It uses first-order optimality conditions of the JKO scheme instead of complex bilevel optimization. JKOnet* can recover potential, interaction, and internal energy components of diffusion processes. The authors provide theoretical analysis and experiments showing JKOnet* outperforms baselines in accuracy, speed, and ability to handle high-dimensional data. They also derive a closed-form solution for linearly parameterized functionals. JKOnet* offers improved computational efficiency and representational capacity compared to existing approaches for modeling diffusion dynamics from population data.

Strengths

- Develops JKOnet*, a method using first-order optimality conditions of the JKO scheme to learn diffusion processes, avoiding bilevel optimization and improving computational efficiency. - Provides theoretical analysis and proofs for JKOnet*, including a closed-form solution for linearly parameterized functionals, backed by comprehensive experiments across various test functions. - Demonstrates improved performance in terms of Wasserstein error and computation time compared to existing methods like JKOnet, especially in high-dimensional settings. - Enables recovery of potential, interaction, and internal energy components of diffusion processes, expanding the model's applicability to more complex systems and improving interpretability.

Weaknesses

- The experimental evaluation is limited to synthetic datasets. Real-world data applications would strengthen the practical relevance of the method. - While the paper discusses limitations, it does not thoroughly explore potential failure cases or boundary conditions where JKOnet* might underperform. - The paper does not provide a comprehensive comparison with other recent approaches in learning diffusion processes beyond JKOnet, which could provide broader context for the method's improvements.

Questions

- The authors demonstrate JKOnet*'s performance on synthetic datasets. How well does the method perform on real-world diffusion processes? Additional evaluations on empirical data would help understand the method's practical applicability. - The paper focuses comparison mainly with JKOnet. How does JKOnet* compare to other recent approaches in learning diffusion processes? - In Section 3.4, the authors discuss different parameterizations. How sensitive is JKOnet* to the choice of neural network architecture for the non-linear parameterization case?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The author discusses limitations in section 5

Reviewer CXdg5/10 · confidence 3/52024-07-18

Summary

This paper studies the problem of learning a diffusion process from samples. It proposes a new scheme based on learning the "causes mismatch" of the process, rather than the "effects mismatch" as in previous works. The new method is significantly more efficient than the schemes from prior works, and works well in practice.

Strengths

The paper is well-written, and the scheme proposed seems to work well in practice on the examples it was tested on. The loss is intuitive, and resembles the score-matching loss from diffusion models, but is the analogous version for arbitrary diffusion processes. Overall, this seems like a paper that people at NeurIPS would be interested in.

Weaknesses

I am not familiar enough with the literature, but it seems surprising to me that this scheme has never been proposed before. In particular, the loss is exactly the score-matching in the case of diffusion models, and there are works [1], [2] that have proposed a similar loss for arbitrary diffusion processes. [1]: https://arxiv.org/abs/2208.09392 [2]: https://arxiv.org/abs/2209.05442

Questions

1) Can you provide a more thorough comparison with prior literature, especially the works I have linked above?

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

N/A

Reviewer fc3q2024-08-08

Acknowledgement of the rebuttal

I thank the authors for their rebuttal. I remain confident of the quality of their paper, suggest the acceptance and keep my score.

Reviewer vFxe2024-08-09

I appreciate the detailed response and additional experiments in the rebuttal, and continue to recommend paper acceptance.

Program Chairsdecision2024-09-25

Decision

Accept (oral)

© 2026 NYSGPT2525 LLC