Summary
This paper explores the mathematical basis for the principle of "guidance" in generative models built out of dynamical transport of measure and provides a mathematical analysis on why certain effects are empirically observed.
They provide this theoretical analysis for a mixture distribution and then test if these results hold in the case of images for classifier and classifier-free guidance.
Strengths
This paper motivates well what its aims are. In addition, the authors provide a suite of experiments ranging from simple synthetic examples to support the main theoretical claims about the evolution of the probability flow ODE under different guidance scales.
Weaknesses
The reviewer finds the paper pretty disorganized, to the point that it is hard to follow the validity of some of the theoretical statements as well as their implication. In particular, I'd like to draw the following comments to the authors in hopes that they can improve these aspects of the paper:
- There are some statements early on that I found confusing, and without theoretical justification. For example, the statement: "In other words, the operation of applying noise to p and the operation of tilting it in the direction of the conditional likelihood do not commute," confuses the reviewer, in the sense that it is not clear why the relation they refer to breaks down. There are proofs in other papers that show that there is a valid transport equation for the classifier guidance setting, e.g. the appendix C in [1]. The question the reviewer thinks the authors should be trying to ask is how to interpret this tilted density (their first unmarked equation).
- The organization of the theorems in the paper makes them a bit hard to follow. The authors introduce Theorems 1 and 2 early on in the motivation of the paper, but don't provide a preliminaries section until 2 pages later that try to introduce some of the distributions under consideration. Following this, there is a section on numerical experiments to motivate these theorems, but then a return to results on the mixture of uniform distributions relevant for the theorems. This organization needs serious work to be compelling. It's pretty hard to follow which aspects of the theoretical results one should be keeping track of to see if the experiments really support them. The paper then ends with a high level sketch of this last proof.
- Certain equations are introduced with no clarification of notation, for example the probability flow ODE (eq 4), nor is it clear where this equation comes from unless you know the literature. It's also unclear why a different formulation of it is included in Lemma 1.
- The reviewer appreciates the efforts of the authors to include results on image generation, however the experimentation is a bit thin and heuristic, only relying on this pullback effect. For example, on the MNIST experiments, how do you quantify this as outside the support of the density?
[1] Ma et. al. (2024) https://arxiv.org/pdf/2401.08740
Questions
Can you please clarify what you mean by this notion of noising a distribution *p* (which doesn't really make sense to me, I think you mean noising the samples e.g. convolving *p* with a Gaussian) and tilting not commuting? I don't fully understand the point still. There is a valid transport equation for *p_t* and therefore also an associated probability flow, so it's a bit unclear to me still by what you mean. See the above paper.
Limitations
The authors have not addressed many limitations of this work, though the reviewer does sympathize with not having easy access to compute or the image models necessary to really do some good experimentation.