Summary
This paper proposes conditional MMD flows with the negative distance kernel for posterior sampling and conditional generative modelling. By controlling the MMD of the conditional distribution using the MMD of the joint distribution, the paper provides a pointwise convergence result. In addition, the paper shows that the proposed particle flow is a Wasserstein gradient flow of a modified MMD functional, and hence provides some theoretical guarantee for [1]. Finally, the paper experiments on several image generation problems and compares with other conditional flow methods.
[1] C. Du, T. Li, T. Pang, S. Yan, and M. Lin. Nonparametric generative modeling with conditional slicedWasserstein flows. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (eds.), Proceedings of the ICML ’23, pp. 8565–8584. PMLR, 2023.
Strengths
1. The paper is well-written and clearly-organized.
2. The paper proves that the proposed particle flow is a Wasserstein gradient flow of an appropriate functional, thus providing a theoretical justification for the empirical method presented by [1].
3. Abundant generated image samples are shown in the experiments.
[1] C. Du, T. Li, T. Pang, S. Yan, and M. Lin. Nonparametric generative modeling with conditional slicedWasserstein flows. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (eds.), Proceedings of the ICML ’23, pp. 8565–8584. PMLR, 2023.
Weaknesses
1. The novelty of the proposed method appears to be limited, since it is mainly the Generative Sliced MMD Flow [1] method applied to conditional generative modelling problems. Additionally, the proof of Theorem 3 partially follows [2].
2. The theoretical comparison with different kernels (Gaussian, Inverse Multiquadric and Laplacian [1]) and discrepancies (KL divergence, W_1 [2] and W_2 [3] distance) in Theorem 2 is insufficient.
3. The numerical results of image generation lack comparison with other methods like Generative Sliced MMD Flow in [1]. It would be better to compare the FID scores for different datasets and various methods like [1], since the proposed method adopts the computational scheme of Generative Sliced MMD Flow. It would be beneficial to compare with Conditional Normalizing Flow in the superresolution experiment and with WPPFlow, SRFlow in the computed tomography experiment.
[1] J. Hertrich, C. Wald, F. Altekrüger, and P. Hagemann. Generative sliced MMD flows with Riesz kernels. arXiv preprint 2305.11463, 2023c
[2] F. Altekrüger, P. Hagemann, and G. Steidl. Conditional generative models are provably robust: pointwise guarantees for Bayesian inverse problems. Transactions on Machine Learning Research, 2023b.
[3] F. Altekrüger and J. Hertrich. WPPNets and WPPFlows: the power of Wasserstein patch priors for superresolution. SIAM Journal on Imaging Sciences, 16(3):1033–1067, 2023.
Questions
1. The paper states that MMD combining with the negative distance kernel results in many additional desirable properties, however it lacks convergence rate or discretization error analysis because “the general analysis of these flows is theoretically challenging”. Regarding this problem, what is the advantage of MMD over other discrepancies like Kullback–Leibler divergence or the Wasserstein distance especially for conditional generative modelling problems?
2. Is it possible to provide a discretization error analysis between discrete MMD flow and the original continuous MMD flow?
Rating
5: marginally below the acceptance threshold
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.