This paper presents a new gradient flow dissipation geometry over non-negative and probability measures. This is motivated by a principled construction that combines the unbalanced optimal transport and interaction forces modeled by reproducing kernels. Using a precise connection between the Hellinger geometry and the maximum mean discrepancy (MMD), we propose the interaction-force transport (IFT) gradient flows and its spherical variant via an infimal convolution of the Wasserstein and spherical MMD tensors. We then develop a particle-based optimization algorithm based on the JKO-splitting scheme of the mass-preserving spherical IFT gradient flows. Finally, we provide both theoretical global exponential convergence guarantees and improved empirical simulation results for applying the IFT gradient flows to the sampling task of MMD-minimization. Furthermore, we prove that the spherical IFT gradient flow enjoys the best of both worlds by providing the global exponential convergence guarantee for both the MMD and KL energy.
Paper
Similar papers
Peer review
Summary
The paper proposes a gradient flow in combined Wasserstein-MMD geometry w.r.t. certain functionals. The authors primarily consider MMD squared functional, but also have some theoretical results regarding KL divergence functional. The work is more about theory: the authors are concerned about some mathematical properties of their proposed flows and convergence analysis, and have only toy 2D illustrative experiments.
Strengths
* Overall, the considered topic is quite interesting. The theory of Gradient Flows in different geometries is an emergent field which is at the intersection of Machine (Deep) learning and mathematics. This theory is full of remarkable, non-trivial, theoretical results. Transmitting all of this mathematical beauty into practical algorithms is praiseworthy. * The paper has some interesting theoretical results and statements.
Weaknesses
* (A) At first, I found the manuscript to be a bit difficult to read, especially section 2. A lot of specific mathematical terms were used, e.g., “Onsager operator”; tangent/cotangent spaces and metric tensor of (probability) measures space. A lot of relationships between these specific objects were mentioned, e.g., $\mathbb{K} = \mathbb{G}^{-1}$; formulation of gradient flow through the Onsager operator (eq. 2). I think that in order to make the text more accessible for those who are not a specialist in geometry of (probability) measure spaces, it should be either simplified, or all necessary theoretical introductions should be done, e.g., in the appendix. * (B) I am not fully satisfied with the structure of the text. In particular: * Why Remark 3.4 and Corollary 3.5. (some properties of pure IFT gradflow which was introduced much earlier) are located right after technical Theorem 3.3 (Lojasiewicz inequalities)? I think it is better to place these statements right after Remark 3.1. * For me, it is a bit strange that the paper develops theory of spherical IFT gradflows (Section 3.1.), while the only practically considered case (where the driving functional for the gradflow is MMD) does not require this theory (Theorem 3.6) because spherical MMD (IFT) coincides with conventional MMD (IFT) flow. May be more emphasize (inculuding practical evaluations) should be put on KL-driving gradflows, where the sphericity matters. * (C) (lines 98-99). Machine learning applications of Wasserstein gradient flows: some missed links: [1-6] * (D) To be honest, I am a bit skeptical about the pure MMD gradient flow (in the MMD geometry) - which is denoted as (MMD-MMD-GF) in Theorem 3.6, and, correspondingly, my skepticism extends to MMD gradflow in the joined Wasserstein-MMD geometry - (MMD-IFT-GF) in Theorem 3.6. At first, (MMD-MMD-GF) was considered in literature, e.g., [7] (not cited!) - see Section 3.1, Case 1 of their paper. And it was noted that such pure MMD-MMD flow is undesirable in practice, exactly because it “teleports” mass between initial and target distribution (note that the solution to MMD-MMD is just interpolation between distributions, as noted by Theorem 3.6). As I understand, the idea of the paper under consideration is that by considering joined Wasserstein-MMD geometry (MMD-IFT flow) one can leverage this problem. However, in the paper, I didn’t find sufficient evidence that it is the case. In particular, the proposed practical optimization procedure includes solving proximal MMD minimizing-movement step, eq. 16. For $F = \text{MMD}$ it boils down to eq. (18), which is MMD barycenter problem. It is known that MMD barycenter problem has solution, see [8, proposition 2], - it is just a mixture of input distributions. Therefore, if we fairly solve MMD minimizing movement step, the resulting $\mu^{\ell + 1}$ will mix $\mu^{\ell + \frac{1}{2}}$ and target $\pi$, i.e., the teleporting of mass will occur. * (E) The practical validation of the method is rather weak. Only a couple of 2D experiments with Gaussians/Mixture of Gaussians. Moreover, I didn’t find that the proposed method performs better than the alternatives. May be, according to some metrics it is indeed the case, but the visual performance of the method is somewhat disappointing. As I understand from Figures 3, 4 and gifs provided in the supplementary, the method leaves a considerable number of points far from the support of target distribution. The alternatives, even vanilla MMD flow, are better in terms of this characteristic. [1] Gao et. al., Deep Generative Learning via Variational Gradient Flow, ICML’2019 [2] Gao et. al, Deep Generative Learning via Euler Particle Transport, MSML’2021 [3] Mokrov et. al., Large-Scale Wasserstein Gradient Flows, NeurIPS’2021 [4] Alvarez-Melis et. al., Optimizing Functionals on the Space of Probabilities with Input Convex Neural Networks, TMLR’2022 [5] Bunne et. al., Proximal Optimal Transport Modeling of Population Dynamics, AISTATS’2022 [6] Fan et. al., Variational Wasserstein gradient flow, ICML’2022 [7] Mroueh et. al., Sobolev Descent, AISTATS’2019 [8] Cohen et. al., Estimating Barycenters of Measures in High Dimensions
Questions
* (a) What are the properties of inverse operator $\mathcal{K}^1$. In particular, why is it linear on its arguments? * (b) The MMD minimizing movement step (eq. 16 and eq. 18) is being solved inexactly in practice, e.g., only weights of particles are optimized, while the exact solution is the mixture of source and target distributions. Moreover, even eq. 18 is substituted with a single step of project GD. What is the reason? How do such approximations affect theoretical and practical properties of the proposed method? * (c) In the appendix, proofs. Why does the equality hold: $\langle \frac{\delta F}{\delta \mu}[\mu], \mathcal{K}^{-1}\frac{\delta F}{\delta \mu}[\mu] \rangle_{L_\mu^2} = \Vert \frac{\delta F}{\delta \mu}[\mu] \Vert^2_{\mathcal{H}}$?
Rating
4
Confidence
3
Soundness
3
Presentation
2
Contribution
2
Limitations
The limitations were addressed correctly
End of discussion period approaching
Dear reviewer, Thank you for your feedback on our manuscript. We have carefully considered your comments and suggestions and have made the revisions. We have also included new experiments and done our best to answer your questions. The rebuttal took a tremendous amount of effort and we want to make sure it has been read. As the discussion period will be closed soon, we kindly ask for your feedback on the rebuttal. Have we addressed your concerns? Is there anything else we can improve? Thank you again for your time and effort. Authors
Summary
This manuscript proposes a new gradient flow over probability and non-negative measures termed the interaction-force transport (IFT). The flow is based on the inf-convolution of the Wasserstein Riemannian metric tensor and the spherical maximum mean discrepancy (MMD) Riemannian metric tensor. The authors provide a number of theoretical results related to their proposal: global exponential convergence guarantees for the MMD and Kullback-Leibler divergence energies. While convergence results are available for the KL divergence energy, not much is known theoretically in the context of the MMD energy. The authors then introduce a particle gradient descent algorithm for IFT, composed of two steps gradient flows (one to update particle locations via the Wasserstein step and one to update particle weights via the MMD step), and show two proof-of-concept examples that validate the developed theory.
Strengths
1. The paper is well-written and well-organized. Despite heavy notation throughout, the authors do a nice job introducing the notation and staying consistent throughout the manuscript. The literature review on related method is also quite thorough. 2. In my opinion, this is a significant theoretical contribution to the machine/statistical learning literature. The topic is timely and will appeal to a broad NeurIPS audience. 3. While the paper is theoretical in nature, I appreciate that the authors provide an implementation of the proposed gradient flows. The presented algorithm is easy to understand ensuring reproducibility. 4. The contribution of this manuscript is original in many aspects: the new definition of gradient flows, the proofs of global convergence for the MMD and KL energies, and comparison to previously defined methods for the MMD flow that required a heuristic noise injection step.
Weaknesses
I did not identify many weaknesses in this work. While I understand that the contributions are theoretical in nature, I wish the authors would have presented more examples/comparisons to the work of Arbel et al. [2019] in an appendix. The presented examples are sufficient as proof of concept.
Questions
Could the authors comment on the roughness of the standard deviation bands in Figure 2 for the proposed method? Also, while this limitation is already mentioned briefly in the discussion, I feel that a bit more could be said about scalability of the presented algorithm. This does not necessarily have to be included in the main body of the paper, but rather in the appendix where the algorithm is presented.
Rating
8
Confidence
2
Soundness
4
Presentation
4
Contribution
4
Limitations
Limitations were adequately addressed.
After considering all of the reviews and the authors' rebuttal, I am inclined to keep my rating as is. I appreciate the additional experiments provided in the PDF file as part of the rebuttal.
Thank you for reading the rebuttal
Dear reviewer, Thank you for taking the time to read and respond to our rebuttal! Your feedback has helped improve our manuscript. Authors
Summary
This paper proposes a novel gradient flow geometry – interaction-force transport (IFT). It is theoretically shown that IFT gradient flow has global exponential convergence guarantees both for MMD and KL energies. The authors propose an algorithm based on the JKO-splitting scheme and test it on examples with 2D Gaussians and Gaussian mixture.
Strengths
The paper provides the proof of the exponential convergence guarantees for IFT gradient flow with MMD energy.
Weaknesses
I am not convinced that the established theoretical results and provided experimental justifications are sufficiently significant. The main theoretical contribution of the current paper is the proof of exponential convergence guarantees for their IFT gradient flow both with MMD and KL divergence energy. (Actually, the proofs of these results do not seem to be very impressive, e.g., the proof of Proposition 3.8 immediately follows from two well-known facts. Anyway, it is not my major concern.) The authors state that the established convergence guarantees (especially, for the MMD energy) are the main motivation for considering the IFT gradient flow. However, for the KL case, it is not a surprising property since even ordinary Wasserstein flow with KL divergence energy has the same exponential convergence guarantees. For the MMD case, the results are quite novel, although there exist several other works which prove some convergence properties of flows with MMD energy (Arbel et al., 2019). Thus, I am wondering, are the provided proofs of exponential convergence rates actually important for the practical use of the designed algorithm? The empirical evaluation of the algorithms seems to be very limited. The algorithm is tested only in low-dimensional experiments using 2D Gaussians and Gaussian Mixtures which immediately raises questions regarding the scalability of the approach. Besides, the authors compare their approach only with Wasserstein flows with MMD energy (with or w/o noise injection) (Arbel et al., 2019, Korba et al., 2021). However, it is important to see how the algorithm behaves in comparison to flows with KL divergence energy as well. Overall, my main concerns are related to the questionable significance of paper results. The proof of exponential convergence rates for their IFT gradient flow solely does not seem to be a significant contribution. Meanwhile, the experimental evaluation needs to be considerably enhanced. *Minor*: - line 218: 'expontial' - typo
Questions
Does your approach have some practical use cases? I suggest the authors to improve the experimental part of their paper by - including experiments in dimensions larger than $d=2$ - performing comparisons with flows using KL divergence as the energy, e.g., with (Yan et al., 2023, Lu et al., 2019) which are cited in the paper I am open to adjusting my score if the authors address these suggestions. **References.** M. Arbel, A. Korba, A. Salim, and A. Gretton. Maximum Mean Discrepancy Gradient Flow. arXiv:1906.04370, Dec. 2019. A. Korba, P.-C. Aubin-Frankowski, S. Majewski, and P. Ablin. Kernel Stein Discrepancy Descent. In Proceedings of the 38th International Conference on Machine Learning, pages 5719–5730. PMLR, July 2021. Y. Yan, K. Wang, and P. Rigollet. Learning Gaussian Mixtures Using the Wasserstein-Fisher-Rao Gradient Flow. arXiv:2301.01766, Jan. 2023. Y. Lu, J. Lu, and J. Nolen. Accelerating Langevin Sampling with Birth-death. ArXiv, May 2019.
Rating
5
Confidence
2
Soundness
2
Presentation
3
Contribution
2
Limitations
The authors have addressed the limitations of their approach in the discussion section.
End of discussion period approaching
Dear reviewer, Thank you for your feedback on our manuscript. We have carefully considered your comments and suggestions and have made the revisions. Per your suggestions, we have also included new experiments and done our best to answer your questions. As the discussion period will be closed soon, we wish to ask for your feedback on the rebuttal kindly. Have we addressed your concerns? Is there anything else we can improve? Thank you again for your time and effort. Authors
I thank the authors for their answers to my questions and concerns. First, I appreciate that you conduct a moderate-dimensional (d=100) experiment with mixture of 3 Gaussians as a target and include the comparison with some of the requested approaches. I expected that you will also provide some figures visualizing the obtained results (although, I understand that it might be quite tricky for d>2). Second, as you explain, the practical implementation of KL energy with the proposed IFT flow (requested by me and other reviewer) is out of the scope of this paper. From my point of view, you should somehow state this directly in the paper, otherwise, the existence of the whole section of the paper related to this case seems to be confusing from my point of view. Third, I am still not sure that the provided theoretical results with only moderate-dimensional experiments (up to d=100) with Gaussians are solid enough to be published. I see that the limited experimental evaluation ('proof-of-concept' type of experiments) was noted by other reviewers too. Respecting the time spent by the authors on running the experiments, I *adjust my score* accordingly. Meanwhile, I am looking forward for further discussion with other reviewers and Area Chairs.
Thank you for responding to our rebuttal
Dear reviewer, Thank you for considering our rebuttal. We appreciate your feedback and that you have adjusted your score accordingly. We agree and will indeed provide a detailed explanation of the sampled-based (MMD energy) vs score-based (KL energy), as outlined in our rebuttal text, especially around Sec 3.3. We respect the reviewer's third point. Indeed, our experiments may be "proof of concept". Our hope is to propose the IFT "gradient structure" in this paper (e.g. K_IFT in eq (7) ) and study its properties (e.g. Theorem 3.6). In view of the already sizable literature on MMD flows started by Arbel et al. (2019), we believe that the proposed gradient structure is a significant contribution and will generate useful mathematical insights for ML researchers working on related topics. From the technical perspective of gradient flows, the discovery of a new gradient structure is already quite non-trivial. That is our intention in this paper. However, we perfectly respect that the reviewer may have a different perspective based on their expertise. We will do our best in the next revision to make our insight useful for a wider audience. Thanks again, Authors
Summary
This paper proposes a novel gradient flow geometry (IFT), based on the infimal convolution of the Wasserstein tensor with the MMD tensor. For this geometry, the authors show global exponential convergence guarantees for both MMD and KL energies. They then develop an algorithm for the IFT gradient flow and test it on an MMD inference task, showing empirically that it avoids mode collapse as Arbel et al. does.
Strengths
1. The exposition clarity is excellent. The authors do a great job of positioning their work relative to existing gradient flow works. 2. In introducing a novel gradient flow geometry and showing favorable convergence characteristics, the work has good potential to inspire follow-on works.
Weaknesses
1. The experiments run are relatively simple and low-dimensional, and it is not clear how practical the method would be for more realistic application scenarios. 2. There is no example comparison of behavior for KL, which would have been nice to see.
Questions
1. Looking at the numerical example, have you tried any heuristic approaches, e.g. some sort of branching, for eliminating and repopulating particles when weights get very low on certain particles? It seems such behavior might improve performance empirically.
Rating
6
Confidence
2
Soundness
3
Presentation
4
Contribution
4
Limitations
Limitations are well-acknowledged by the authors.
End of discussion period approaching
Dear reviewer, Thank you for your feedback on our manuscript. We have carefully considered your comments and suggestions and have made the revisions. We have also included new experiments and done our best to answer your questions. As the discussion period will be closed soon, we wish to ask for your feedback on the rebuttal kindly. Have we addressed your concerns? Is there anything else we can improve? Thank you again for your time and effort! Authors
Thank you & keeping score as is
Dear authors, Thank you for the clear and extensive response to my review. I think I will maintain my score as is, remaining slightly positive on the paper, given the proof-of-concept experiments and perhaps limited immediate impact. Reviewer
Thank you for reading our rebuttal
Dear reviewer, Thank you for your feedback and for taking the time to read the rebuttal. Best regards, Authors
Thanks to the authors
I thank the authors for the answers they provided and appreciate the fairly expressed attitude towards my review. Some comments: 1. Indeed, I missed that your method supports the weights on par with the particles itself, and this is the reason of my wrong evaluation of your 2D Gaussian experiments. My bad, my carelessness. It seems that I was a bit biased due to typical particle flows I know (MMD flow, KSD flow, SVGD flow) which do not introduce this additional complexity with weights. 2. I appreciate the additional 100D Gaussian $\rightarrow$ mixture of 3 Gaussians experiment. It strengthens the work. 3. **The reviewer claimed our theory is "not needed"** - I never said that anywhere in my review. The only my "not needed" was about ethics review. Regarding the theory, I just wanted to encourage the discussion on the practical aspects of the flows different from MMD (IFT). 4. Which particular work do you mean by **[Y. Mroueh and M. Rigotti]**? Also, I am a little confused as to why the authors refuse to cite a related work [7], which I pointed out in my review. 5. For sure, MMD-MMD-GF is not your focus. My point was that MMD-IFT-GF (the only practically evaluated flow you propose) inherits some undesirable properties of the pure MMD-MMD flow. I mentioned the teleporting of mass problem noticed in the old paper I cited. And I just claimed that this teleporting of mass problem also appears in your case (MMD-IFT-GF) - theoretically - when solving (MMD minimising movement step) - eq. 16. This is because the MMD minimizing movement step (as you noticed) boils down to MMD barycenter problem, which has a known solution (see [8, proposition 2]). And this known solution is just a mixture of source and target samples. Maybe this teleporting of mass phenomenon is indeed clear from your PDE/ODE characterization in Thm 3.6, but anyway this phenomenon worth to be explicitly mentioned in the paper 6. **In optimization, it is often desirable and significantly faster to perform inexact iterations to be more computationally efficient.** In general, I agree with this statement. However in your case solving the MMD barycenter problem exactly seems to be faster, because, as [8, proposition 2] notices, the solution is just the mixture of particles. In conclusion, I thank the authors one more time and raise my score.
Thank you for acknowledging the major misunderstanding. We have now addressed the new points.
We thank the reviewer for reading our rebuttal. However, the 6 points raised by the reviewer appear to digress from the main point of the rebuttal and do not justify the rejection assessment. We now point-by-point address the reviewer's comments. Due to the lateness of the comments, we try our best to be thorough. ### Points 1 and 2 Thank you for acknowledging the major misunderstanding and acknowledging our new experiments. We believe those issues are now resolved. ### Point 3 > Under item (B) of "Weaknesses", it was stated that "the only practically considered case (where the driving functional for the gradflow is MMD) **does not require this theory** (Theorem 3.6)". First, we apologize for wrongly writing "required" as "needed", though we believe the meaning is the same. This is what our "not needed" comment refers to. Does the reviewer's claim "not require" does not imply "not need", or does the reviewer still stand by this assessment? In any case, we have already addressed this (non)-issue in the rebuttal, and we believe this point is now resolved. ### Point 4 [Y. Mroueh and M. Rigotti]: Unbalanced sobolev descent. Advances in Neural Information Processing Systems. 2020;33:17034-43. > I am a little confused as to why the authors refuse to cite a related work [7] We have clearly stated in the rebuttal that the newer paper above contains the old framework but additionally newer methodologies (e.g. Kernel-Sobolev-Fisher discrepancy, which is more general) and results. Proper scholarly practice is to avoid block citations of many similar papers, when the relevant line of work has already been covered. We also need more time to look into the content of the 7 papers [1-7] the reviewer suggested we cite. [Y. Mroueh and M. Rigotti] appears to be more recent and comprehensive to the best of our knowledge. We hope our reason is clear. In any case, we also did not refuse to cite [7], we simply stated that we have covered kernel Sobolev descent, and it should not be used as an argument to undermine our contribution. Furthermore, none of those papers contain the contributions of our paper, so we again wish to emphasize that comments digress from the main point of the rebuttal and the rejection assessment is not justified here. ### Point 5 > My point was that MMD-IFT-GF ... inherits some undesirable properties of the pure MMD-MMD flow. We still do not see any mathematical justification for this "inheritance". The comments kept mentioning the MMD steps, but the IFT has a Wasserstein step with diffusion. So the comment is not sound. More mathematical analysis and evidence are needed to support such a claim. > anyway this phenomenon worth to be explicitly mentioned in the paper We agree. We have already done this by giving the precise mathematical formulation of the flow solution: see the second formula in Thm 3.6. This precise statement is already explicit. Furthermore, the solution of (MMD-IFT-GF) has not been studied and is not simply a mixture. We are open to adding more plain English sentences to the presentation if it helps non-experts understand the results better. But again, the reviewer digresses and this is a minor presentation (non-)issue. > known solution is just a mixture of source and target samples Mathematically, this is not rigorous. We do not assume the target distribution to be discrete, the mixture is infinite-dimensional and not simple to implement. The goal of IFT or Arbel et al.'s work is to find a gradient-based algorithm to generate the path $\mu_t$. Again, we do not see this to be a mathematical reason to undermine IFT. ### Point 6 We agree that there can be more discussion and future work on how to implement the MMD step. We already discussed with great detail. We will improve the presentation. > solving the MMD barycenter problem exactly seems to be faster We have tried both in practice. What the reviewer described is not the case. "Faster" for what? We will expand on this in the corresponding section. > because, as [8, proposition 2] notices, the solution is just But our goal is not to solve the MMD Barycenter sub-problem -- it is just a subroutine in the JKO splitting scheme. The goal is to generate samples to construct the path $\mu_t$. The logic in the reviewer's comment is flawed: why not just take samples from $\pi$ directly? Why do we bother using methods such as that from [Arbel et al.]? Furthermore, [8, proposition 2] simply says the solution to the sub-step is a reweighting, that is precisely what we implemented. One must optimize the reweighting coefficients $\beta_p$ in [8]. The reviewer's comments seem to have trivialized the this. Again, zooming out to the big picture, this numerical detail of implementation, for which we did not hide anything, does not seem to be a sound case for rejecting our contribution. ### Conclusion In summary, we appreciate the reviewer's time and effort. However, those comments do not justify the rejection assessment.
Decision
Accept (poster)