Summary
Building upon the concept of probability flows, the paper establishes a clear link between diffusion processes and optimal transport. It first shows that the continuous probability flow corresponds to a Monge map for every finite-length time interval. It then states and proves analogous statements in the context of continuous-time discrete space (CTDS) models. These results follow the intuition that diffusion should be asymmetric to avoid _mutual flows_ between states. Finally, the authors design a practical sampling method that substantially reduces the variance of end positions of particles, given their initial status.
Strengths
The paper presents a well-executed analysis of diffusion processes from the perspective of probability flows. It modifies a previous result by Khrulkov et al. [29] to make it applicable to the case of Brownian motion with initial distribution given by Eq. 17, i.e., that of a delta train induced by samples.
To analyze the discrete case, the authors revise the (reversed) heat diffusion process by preventing moves to states with lower probability. This procedure, reminiscent of well-known sampling techniques, leads to easy-to-analyze forward (Eq. 27) and backward transition rates. The authors show that the resulting process displaces mass optimally with respect to the L1 cost.
Weaknesses
The notation used throughout the paper is at times hard to parse: Several similar symbols are used concurrently but there is no concise overview of their meaning (and differences). This makes the presentation less intuitive and requires more effort from the reader. In particular, the interpretation of “infinitesimal transport” would benefit from a revision aimed at better disambiguating between processes $\hat{Y}$ and $\tilde{Y}$, and between generators $A$ and $\hat{A}$. Also, the expanded state notation (i.e., $i_1, …, i_K$) is never needed in the main text and could be dropped altogether.
It would also be good to include additional derivations (in the appendix), such as the one of the reverse-time generator $R$ (eq. 29), which are not entirely trivial.
**Experiments**
The authors include several experiments on toy datasets that showcase the effectiveness of the proposed sampling strategy in reducing the variance. They nicely illustrate the theoretical statement which links the process generated by $Q$ to an optimal transport map. However, they do not describe any practical scenarios in which this feature could be necessary or desirable. In particular, reduced variance in sampling seems to imply that slight biases in the initial distribution of particles could lead to unfaithful reconstructions of the other marginal, i.e., by creating “holes” that the optimal transport map fails to fill.
I am also skeptical of the choice of experiments. Why is the current selection restricted to synthetic and low-dimensional settings? Did you run out of time or is the method's applicability limited?
**Miscellaneous**
To ease the description of sampling processes it would be helpful to use the differential (or discrete increment) of time flowing backward, e.g., in Eq. 2.
Further, some **typos**:
- In Table 1, given that the MMD value for the SDDM simulation on the checkerboard dataset has the wrong sign.
- In line 72: $(x_{t+t})$.
- In Eq. 19: missing a ½ factor?
- In the Discussion section.
Questions
What are possible practical applications of the proposed sampling method?
What are possible alternative approximations of the quantity $Q_t$, which is likely the cause of the inferior quality of reconstructed marginals?
Why do points in Figure 2 appear to form rectangles, e.g., in the 2spirals, pinwheel, and swissroll datasets?
Even though the infinite horizon case is outside the scope of the paper, it would be interesting to briefly discuss why the arguments presented in the paper break down in that case. Is this somehow related to mixing times of Markov chains? Additional insight, e.g., in the discussion, could provide a valuable starting point for further research.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The practical relevance of the proposed sampling framework is not discussed. It is unclear whether the reduced variance achieved by it could be of use in real-world scenarios and justify the (admittedly) inferior quality of marginal reconstruction quality of DPF.