Energy-Based Sliced Wasserstein Distance

The sliced Wasserstein (SW) distance has been widely recognized as a statistically effective and computationally efficient metric between two probability measures. A key component of the SW distance is the slicing distribution. There are two existing approaches for choosing this distribution. The first approach is using a fixed prior distribution. The second approach is optimizing for the best distribution which belongs to a parametric family of distributions and can maximize the expected distance. However, both approaches have their limitations. A fixed prior distribution is non-informative in terms of highlighting projecting directions that can discriminate two general probability measures. Doing optimization for the best distribution is often expensive and unstable. Moreover, designing the parametric family of the candidate distribution could be easily misspecified. To address the issues, we propose to design the slicing distribution as an energy-based distribution that is parameter-free and has the density proportional to an energy function of the projected one-dimensional Wasserstein distance. We then derive a novel sliced Wasserstein metric, energy-based sliced Waserstein (EBSW) distance, and investigate its topological, statistical, and computational properties via importance sampling, sampling importance resampling, and Markov Chain methods. Finally, we conduct experiments on point-cloud gradient flow, color transfer, and point-cloud reconstruction to show the favorable performance of the EBSW.

Paper

References (57)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer krQV6/10 · confidence 4/52023-07-03

Summary

This paper introduces a new distribution of slices of the Sliced-Wasserstein distance based on the energy-based model framework. The authors study the theoretical properties and conduct numerical experiments on EBSW.

Strengths

- The study of the idea is organized and clear. - Theoretical quantities of interest are derived. - Various sampling methods are provided.

Weaknesses

- The method is not compared with the vanilla Distributional Sliced-Wasserstein method which should be the baseline to beat. The fact that only the vMF-DSW method is benchmarked is a weakness, as the learned distribution is unimodal, which prevents the model to be too expressive. - I think it would be good to have a theorem studying the following sample complexity, that is more of interest for practical reasons: $ \mathbb{E}[EBSW_p(\mu_n,\nu_n;f) - EBSW_p(\mu,\nu;f)]$ - Experiments are a bit light, I think a generative modeling experiment would be nice to have as this is one of the main applications of the SW distances.

Questions

- Can the authors explain why experimental results are not compared to the DSW distance? - It appears to me that EBSW falls into the family of adaptive Sliced-Wasserstein distances ( https://arxiv.org/pdf/2206.03230.pdf , which doesn't seem to be cited in the related work section). Can the authors comment on that and whether this framework can help studying this distance?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

None. I am willing to discuss with the authors and accordingly modify my score.

Authorsrebuttal2023-08-22

Thank You

Dear Reviewer krQV, Thanks for your reviews of our paper. Since the discussion period between the authors and the reviewers was already over and we have not heard from you during this period, we would be grateful if the reviewer could let us know if all your questions are addressed to some extent. If you are satisfied with our answers, we hope that the reviewer will consider adjusting your score. Best, The Authors

Reviewer krQV2023-08-22

Answer to rebuttal

Dear reviewers, I would like to apologize for the late answer and to thank you for the rebuttal. I am happy with the comparison with DSW that I hope will be added to the paper. I have the feeling that a little bit more effort should have been put into proving the triangle inequality for EBSW and the two-sided sample complexity even though I don't know how hard it is to prove. I am however raising my score from 5 to 6 as I think the authors did a good rebuttal.

Reviewer vaua6/10 · confidence 4/52023-07-06

Summary

This paper proposes an extension to Sliced Wasserstein Distance (SW), an approach for measuring distances between distributions by computing the average of the energy of the 1-d Wasserstein distances between 1-d projections. The authors argue that moving towards non-uniform ways of sampling the projections is key, and subsequently show that the benchmark approach currently that does so (Distributional sliced Wasserstein) adds another optimization over a parametric distribution family, which may not yield stable results and is more computationally expensive. Subsequently, they propose a simple non-parametric extension to SW called Energy based SW (EBSW) that achieves the same objective of sampling non-uniformly, i.e. sampling projections with larger Wasserstein distance with a greater probability, by considering energy functions $f$ which are monotonically increasing w.r.t the sliced 1-d Wasserstein distance itself. Theoretical properties are highlighted, including the convergence properties of their proposed metric. Next, the authors propose a range of Monte Carlo estimation methods to approximate EBSW, including deriving the necessary gradient expressions to be used in some of their subsequent applications. Lastly, over multiple experiments ranging from point-cloud gradient flows to deep point-cloud reconstruction, the authors demonstrate that EBSW is faster (Importance Sampling approach), and yields distributions closer to the ground truth in most cases than other SW-based benchmarks.

Strengths

I really liked reading the paper and going through the various findings in it. The following are what I consider to be some of the main strengths of this work. (i) The paper is really well-written. All sections were intuitive to read and easy to understand for me, and this also applied to the Appendices, including the proofs of the Theorems and Propositions. (ii) The contributions of the paper are very clear and supported by the theoretical and empirical results. All proposed benefits of the approach have been verified by the authors in theory and practice. (iii) The experiments are well motivated and relatively exhaustive. The proposed approach is tested in a wide range of problems and the results convincingly showcase the benefits of EBSW. (iv) The background section is well organized and is an effective introduction to SW-based metrics in general.

Weaknesses

The paper is overall interesting and well-written. However, there are a few points noted below that can improve the work further. 1. I feel that the theoretical results can be given a bit more perspective w.r.t. the other SW-based metrics. Overall, I felt the theoretical results in Theorems 1, 2 and Propositions 1 and 2 are relatively intuitive to show. Some more specific insights on EBSW could be interesting (perhaps with constrained energy functions, i.e. Lipschitz constrained $f$). 2. I understand the overall direction for moving towards measures that look disproportionately at projections which have larger distances (lines 37-38). However, I feel that a few more words to elaborate on intuitively why one must move towards non-uniform distributions in the Sliced Wasserstein setting could be useful. 3. One of the main questions I have is the role of $p$, which is fixed for all experiments in this work. I feel that the authors can elaborate a bit on whether trying larger values of $p$ can achieve a similar goal to what EBSW-e achieves (More details in questions). 4. The result tables are convincing and informative w.r.t. the tested methods. However, in some of the visual depictions, I found it hard to see any significant differences between EBSW and some other variants (even in the Appendices). Mainly, the color transfer experiments were a bit hard to discern for me in terms of relative performance

Questions

In addition to the list of weaknesses above, below are some questions and general suggestions for the authors to address: 1. I noticed that the authors fix the value of $p$ to 2 for all experiments. I wonder if increasing $p$ can achieve a similar effect to using a monotonically increasing energy function (such as the exponential distribution in this work). Because intuitively, it seems to me that increasing $p$ eventually also gradually assigns more importance to the projections with values near to the maximum. I’m mainly asking this because increasing $p$ still falls under SW-p, and thus could perhaps have similar sample complexity and computational load. Could you please provide a theoretical explanation on how changing $p$ fundamentally differs from EBSW with an increasing energy function? 2. For Figure 2, the authors state that the gradient flows for EBSW are smoother than other approaches. Perhaps describing some visual cues to elaborate on that point can make it easier to follow. 3. Although four different ways of estimating the EBSW variants are proposed, I only see the IS-EBSW being reported in the main paper. Furthermore, having looked at the results in the appendix, it seems in almost all cases, the variants have roughly similar performance, and IS-EBSW seems to lead to better results in most cases. Furthermore, it also seems that the time complexity for all algorithms are similar, also seeing Table 3 in C.1, it seems that in most cases IS-EBSW is faster as well. Thus, my question is: in discussing these other approaches, are there benefits to the other proposed MCMC variants? 4. Are there any cases where the choice of an exponential energy function $f$ doesn't work? It would be intuitive to me that, considering a more general family of exponentials in $e^{(\lambda x)}$, the optimal "width" of the slicing distribution as set by $\lambda$ may be different for different cases. For instance, can there be any cases where the underlying distribution enforces a much larger or smaller $\lambda$ than one? Or is $\lambda=1$ in some sense universal? 5. As a follow-up to the previous question, I would assume that for the optimization results in Tables 1 and 2, as the generated distribution edges closer to the ground truth distribution, would it be profitable to increase $\lambda$ higher than one? As then the distances would all get very small, and potentially max-SW would start to look more appealing. 6. The theoretical results are interesting and important, however, could the authors put into perspective how the theoretical properties of EBSW compare to other SW based measures? It seems to me that both Theorems 1 and 2 may hold for SW based measures (at least SW, not sure about max-SW and DSW). Similarly, for Proposition 2 it seems from the proof that it would hold for both SW and max-SW (not sure about DSW). Additional theoretical results in this direction, or at the least a rough intuitive discussion on how these results would compare for the other SW counterparts, would be useful. This would put more into perspective how EBSW's theoretical properties compare to the other SW benchmarks.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Some limitations have been discussed, but I feel there is scope to discuss more limitations of the proposed approach.

Reviewer dYij6/10 · confidence 3/52023-07-07

Summary

This paper has proposed an energy-based slicing distribution that maps original distributions into one-dimesion space to compute the Wasserstein distance.

Strengths

This paper has proposed an energy-based slicing distribution that maps original distributions into one-dimesion space to compute the Wasserstein distance. The proposed method shows great advantages over other slicing distribution like sliced Wasserstein, Distributional sliced Wasserstein, and Max sliced Wasserstein. The authors also explore different sampling methods to approximate the value of the EBSW distance.

Weaknesses

1. An introduction about optim transport should be included in the Background section. 2. Put a title for each row in Figure 2 could help reader easier to compare different SW methods.

Questions

1. In definitition 1, does $\theta$ represent parameters in the slicing distribution or just the transport function? If $\theta$ are parameters, why could the model be called as parameter-free? 2. In general, the energy function should be defined as $f:(-\inf, \inf)^d \rightarrow [0, \inf]$. Why the author defines it as in line 144? 3. Could the proposed distance measurement be used in generative models, like GANs and EBMs? 4. Is there any comparison with other probability distance measurement methods, like KL divergence, in the experiments?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

This paper has proposed an energy-based slice distribution to compute Wasserstein distance between two distributions. My question mainly lies in the application of energy function and comparison of methods based on other probability measurement methods.

Reviewer QgRs5/10 · confidence 3/52023-07-10

Summary

This paper proposed a new variant of sliced Wasserstein distance that is inspired from Energy-Based Model called Energy Based Wasserstein Distance (EBWD). The proposed method models the energy function of the distance, from which the slices can be sampled. Three sampling techniques are proposed. The paper evaluate the proposed distance calculation using 3 tasks, Point Cloud Gradient flow, color transferring, and Point Cloud reconstruction. The results have shown that the new variant helps the training process convert faster and the transition between two distribution is smoother (in term of Wasserstein distance and Sliced Wasserstein distance getting small faster).

Strengths

The paper has the following strengths - A clear motivation to use the energy function. Similar to SW, the proposed energy-based SW enjoys a non-optimization based computation via sampling using the energy function. - A detailed theoretical analysis of the proposed distance, accompanied by a nice experimental setup to evaluate its performance.

Weaknesses

In general, I like the simple, yet interesting proposal from the paper. However, I also have a few concerns: - It would be reasonable to evaluate EBSW in a more complex task such as image generation, similar to that where DSW is introduced. It seems like only IS-EBSW enjoys the low-computation benefit but other EBSW variants do not. - It is not clear to me why EBSW is better than DSW while DSW finds the most separating directions. Especially the fact that EBSW converges faster.

Questions

Please see comments in Weaknesses

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The paper does not include a Limitation discussion.

Area Chair ncic2023-08-11

Discussion period

Dear reviewers and authors, Thank you very much for your work on this submission and its evaluation. Now that the authors have responded to the reviews, I *strongly encourage* the reviewers to acknowledge the review, to look at other reviews and rebuttals for this submission, and to adjust their scores if needed. Thanks to those that have already done so. Authors have the possibility to reply if further questions are needed, until the 16th. Thank you very much to all, Area Chair

Authorsrebuttal2023-08-15

Reminder on the end of discussion period

Dear reviewers and area chair, First, we would like to express our gratitude again for receiving constructive feedback and questions from the reviewers to improve our paper. We have answered all raised questions in our rebuttal and added new experiments as suggested. Therefore, we would be grateful if the reviewers could give us comments again. Moreover, we are happy to discuss new questions from the reviewers and the area chair. Finally, if the reviewers feel that there are no concerns left, we will be more than happy if the reviewers can increase the assessment score. Best regards, Authors

Reviewer dYij2023-08-19

Thanks for the feedback from the authors. I will remain my rating.

Authorsrebuttal2023-08-19

Response to reviewer

We would like to thank the reviewer for keeping a positive score of 6. We are happy to discuss more if the reviewer still has questions. Best regards, Authors

Reviewer vaua2023-08-21

I thank the authors for their detailed replies to my questions and comments. I am happy to keep my rating. Thanks again!

Authorsrebuttal2023-08-21

Response to Reviewer

We want to thank the reviewer for keeping the score at 6. We will include our discussion in the revision of the paper. Best regards,

Reviewer QgRs2023-08-21

Thank you for the responses!

Thank you for the detailed responses. Please see my additional comments: Q1. Thank you for the additional results on image generation. I notice that the FID of DSW is significantly better than that reported in the original DSW's paper. Also, all the methods have quite similar FIDs, which is difficult to make a conclusion as mentioned in the rebuttal. Q3. It's quite unfair to report the performance of DSW when it does not reach optimality. I think it is ok to report the performance of DSW even if it is better, although it incurs more computation. For example, one can trade-off between different estimation approaches, depending on how much computational budget they have. Currently, the experimental results seem to favor the proposed method also in quantitative metrics as well, which can be misleading. For the other questions, thank you for the clarifications. I will consider all the responses in my final rating of the paper.

Authorsrebuttal2023-08-21

Response to Reviewer

Thank you for your reply, On Q1: The reason the FID score is better since we utilize a stronger backbone for the generative models i.e., ResNet50. As mentioned in the global rebuttal, the main aim of the additional generative modeling experiments is to show the ability to apply widely of EBSW. On Q2: It is worth noting that we do not try to report unfairly for DSW. It is hard to know when DSW will reach its optimality and it could take a lot of computation. Due to the time limitation of the rebuttal, we must focus on the budget-constraint setting where we fix the budget of computation for all baselines. Since there are only 2 hours left of the discussion period, we cannot run additional experiments. However, we will add experiments on letting DSW have a large number of optimization updates in the revision. In the paper, we mainly focus on the computational aspect since EBSW is motivated by the computational limitation of previous variants. We will try to highlight this focus further in the revision. Since the paper is borderline now, we would be grateful if the reviewer could increase the score if all questions are addressed to some extent. We are happy to discuss more if the reviewer is not satisfied with our answers. Best regards,

Reviewer QgRs2023-08-21

Response to the authors!

Thank you for the responses to my questions. It is not my intention to ask for more experiments. The existing evaluation of the paper is overemphasizing the advantages of the proposed method, with limited guidance to its limitations. However, I think the proposed method is interesting and I will consider the responses in its final rating.

Authorsrebuttal2023-08-21

Response to Reviewer

Thank you for your quick reply. We are happy to get the novelty recognition from the reviewer for our proposed method. In addition to the discussed limitation in the global rebuttal, we will try to expand it further in the revision of the paper. Best regards,

© 2026 NYSGPT2525 LLC