Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry

Despite the success of diffusion models (DMs), we still lack a thorough understanding of their latent space. To understand the latent space $\mathbf{x}_t \in \mathcal{X}$, we analyze them from a geometrical perspective. Our approach involves deriving the local latent basis within $\mathcal{X}$ by leveraging the pullback metric associated with their encoding feature maps. Remarkably, our discovered local latent basis enables image editing capabilities by moving $\mathbf{x}_t$, the latent space of DMs, along the basis vector at specific timesteps. We further analyze how the geometric structure of DMs evolves over diffusion timesteps and differs across different text conditions. This confirms the known phenomenon of coarse-to-fine generation, as well as reveals novel insights such as the discrepancy between $\mathbf{x}_t$ across timesteps, the effect of dataset complexity, and the time-varying influence of text prompts. To the best of our knowledge, this paper is the first to present image editing through $\mathbf{x}$-space traversal, editing only once at specific timestep $t$ without any additional training, and providing thorough analyses of the latent structure of DMs. The code to reproduce our experiments can be found at https://github.com/enkeejunior1/Diffusion-Pullback.

Paper

Similar papers

Peer review

Reviewer iEgM6/10 · confidence 4/52023-07-05

Summary

The paper presents an analysis of the latent structure of diffusion models using differential geometry. The authors propose a method to define a geometry in the latent space by pulling back the Euclidean metric from the U-Net bottleneck space *H* via the network encoder. This approach enables the identification of directions of maximum variation in the latent space. The paper also explores the application of the proposed latent structure guidance for image editing. Finally, the evolution of the geometric structure over time steps and its dependence on text conditioning are investigated.

Strengths

- I found the analysis presented in the paper interesting. It both uncovers unknown details about diffusion models (effect of text prompt and complexity of the dataset on the latent space) and confirm some previous observations (e.g., coarse-to-fine behaviour). This exploration can potentially reveal new capabilities of diffusion models, contributing to the advancement of the field. - The paper is technically sound, and the claims made by the authors in sections 4 are supported by experiments. This experimental validation enhances the credibility of the proposed approach.

Weaknesses

- The paper lacks comparisons with other diffusion-based image editing techniques, like [7,18]. Including such comparisons would have provided a more comprehensive evaluation and demonstrated the advantages of the proposed method. - The presentation and clarity of the paper could be improved. For example, the abstract contains too much detail, making it challenging to understand upon initial reading. Also the explanation of the image editing technique could be improved: what is DDIM (section 4)? what is epsilon in Equation 4? Finally, Figure 1 is not sufficiently clear to me, it may hampers the reader's comprehension.

Questions

To make the paper applicable to real-world scenarios right away, the authors should include a comparative analysis with other image editing techniques. **Some potentially interesting references** - Pan, Xingang, et al. "Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold." arXiv preprint arXiv:2305.10973 (2023). - Brooks, Tim, Aleksander Holynski, and Alexei A. Efros. "Instructpix2pix: Learning to follow image editing instructions." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023. - Meng, Chenlin, et al. "Sdedit: Guided image synthesis and editing with stochastic differential equations." arXiv preprint arXiv:2108.01073 (2021). **Typos** - line 20: double citation - line 160 editted - Figure 2 caption: editing

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

The authors have addressed the limitations of their approach.

Reviewer 6jtZ4/10 · confidence 4/52023-07-05

Summary

In this submission, the authors probe the latent space, xt ∈ X, of diffusion models (DMs) from a geometric perspective, utilizing the pullback metric to identify local latent basis in X and corresponding local tangent basis in H. To confirm their findings, they edit images via latent space traversal. The authors provide a two-pronged analysis, investigating the evolution of geometric structure over time and its variation based on text conditioning in Stable Diffusion. Notably, they discovered that in the generative process, the model prioritizes low-frequency components initially, moving to high-frequency details later, and that the model's dependence on text conditions reduces over time. The paper introduces image editing through x-space traversal and to offer comprehensive analyses of the latent structure of DMs, with a specific emphasis on the use of the pullback metric and the SVD of the Jacobian in computing a basis.

Strengths

The paper presents a distinctive idea that provides an alternative method for editing in diffusion models, as well as enhancing comprehension of the latent space dynamics. By utilizing a geometric perspective, the authors make use of the pullback metric to investigate the latent space, offering insights into its structure and operation. The exploration of the evolution of geometric structure over time and its response to various text conditions offers additional insights into the dynamics of the latent space is interesting.

Weaknesses

The first area where the paper could see improvement is in terms of the clarity of its analysis. Given its nature as an analysis paper, it's crucial that the analysis presented is as comprehensible as possible. However, the method and notation used in this work can lead to some confusion. For instance, Section 3, in its current form, may not be as accessible to all readers as it could be, and it could benefit from being revised for clearer communication of the ideas contained therein. Additionally, Fig. 1, which is presumably intended to illustrate key concepts, is perhaps too dense with information. Dividing Fig. 1 into two separate figures could make it easier to digest, enabling a clearer explanation of the approach. A second aspect that could be improved upon is the overall presentation and proofreading of the paper. While the approach is relatively simple, its translation into the written form has not been as clear as one would hope. The text could benefit from a thorough proofreading to ensure that it is not just grammatically correct, but also that it conveys the authors' ideas in a way that is accessible to the wider machine learning community. As it stands, the paper's usefulness to this community may be hindered by its presentation. Lastly, the paper could do more to address the computational implications of its approach. The authors use the power method to approximate the Jacobian, which, while effective, can be computationally costly. It would be beneficial if the authors were more transparent about this fact, allowing readers to fully understand the computational demands of the approach and evaluate whether or not it would be feasible in their own applications. Being upfront about such limitations can help to build a more honest and comprehensive understanding of the paper's methodologies and implications.

Questions

What is the complexity of computing the Jacobian? What is the the actual runtime (in seconds) of the approach compared to other editing methods? I think answering these questions can provide a good context for readers.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

1 poor

Contribution

3 good

Limitations

Societal impact has been properly address but I would like to see a deeper analysis of the computational complexity and runtime of the approach.

Reviewer vuQy5/10 · confidence 5/52023-07-05

Summary

This paper studies the geometry of latent spaces of diffusion models (Dms) using the pullback metric. In the analyses, they mainly examine change of the frequency representations in latent spaces over time and the change of the structure based on text conditioning. After the rebuttal: I checked all reviewer comments and responses. I agree with the other reviewers regarding limited algorithmic novelty of the work. Therefore, I keep my original score.

Strengths

- The paper is well written in general (there are several typos but they can be fixed in proof-reading). - The proposed methods and analyses are interesting. Some of the theoretical results highlight several important properties of diffusion models.

Weaknesses

1. Some of the statements and claims are not clear as pointed in the Questions. 2. The results are given considering the Riemannian geometry of the latent spaces and utilizing the related transformations (e.g. PT) among tangent spaces on the manifolds. However, vanilla DMs do not employ these transformations. Therefore, it is not clear whether these results are for vanilla DMs or the DMs utilizing the proposed transformations. 3. A major claim is that the proposed methods improve effectiveness of the DMs. However, the employed transformations can increase the footprints of DMs.

Questions

It is stated that “To investigate the geometry of the tangent basis, we employ a metric on the Grassmannian manifold.” However, the space could be identified by another manifold as well. Why and how did you define the space by the Grassmannian manifold? It is claimed that “the similarity across tangent spaces allows us to effectively transfer the latent basis from one sample to another through parallel transport”. How this improves effectiveness was not analyzed. In general, how do the proposed methods improve training and inference time? Indeed, the additional transformations can increase training and inference time. Could you please provide an analysis of the footprints?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

2 fair

Presentation

2 fair

Contribution

2 fair

Limitations

Some of the limitations were addressed but potential impacts were not addressed.

Reviewer cKEc7/10 · confidence 3/52023-07-10

Summary

The paper proposes a study on the latent space of diffusion models and on how to manipulate it. It takes advantage of an observation made by previous work [22] on the flatness and semantic structure of the U-Net model used in DDIM and uses pullback metric from the latent space of the U-Net to the space of diffusion to measure some properties of the latter under different conditions. The paper also proposes a method to manipulate the diffusion space through the different time steps, so as to carry out semantically meaningful edits.

Strengths

This paper is one of the first to study the behavior of the space diffusion models. It presents some interesting studies on the behavior of the process during time, showing that the early stages convey higher frequency while the last steps are more concerned with higher frequencies. Another interesting observation is that tangent spaces of different samples tend to be more aligned at T=1, while they diverge toward T=0 (end of the generative process). The paper also shows (even if it had already been observed in 22) that the pullback metric is effective in transferring the shift along the semantically meaningful principal components of the U-Net latent space into the diffusion process, thus resulting in meaningful edits of the generated image, which frequency depends on the time the edit was performed.

Weaknesses

I may have misunderstood or missed some important information, but the method described is not really clear. Specifically, it is not really clear to me how the editing process works (sections 3.3 and 3.4): - In 3.3, the letter v is used to indicate elements of both T_x and T_h, so it is not always clear to which space they are referring. - In general, it is not clear why the idea expressed in 3.3 is useful and where it was adopted. - In eq 4 what is the epsilon function? In general, isn’t the vector toward which to shift selected from T_H (so it should be u) and then transferred to T_x? Another concern is about the generalization of the proposed method to other diffusion techniques, or with other score models (i.e. not UNet). I think that this point needs more discussion.

Questions

I was wondering what would happen if, instead of moving along one of the principal axis of T_H, you use directly the principal axis of T_x. I would also discuss a bit more method 22 in the related work since it seems related to the proposed method. Writing issues: - Check the sentence at rows 74-75 - Row 152: we aims - general grammar check

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

I don’t foresee any particular negative societal impact. A discussion on how the proposed study may generalize to other domains and architectures would be of value.

Reviewer 6jtZ2023-08-15

I have read the rebuttal and thank the authors for their answers. Given the new details provided by authors in terms of computational complexity and runtime of their approach I cannot update my score.

Authorsrebuttal2023-08-15

Thank you for dedicating time to share your thoughts. Firstly, we would like to mention that the main contribution of our paper is the geometric interpretation of the latent space and feature space. To the best of our knowledge, the intricate interplay between semantics and the geometric structure within the latent space of diffusion models remains unexplored. We believe that our paper will significantly enrich future research. \ (Also, please consider the other points written in the global comment.) Moreover, we highlight that our method achieves a comparable complexity, even though efficiency was not our primary focus. \ (We would like to kindly clarify that the 100 seconds mentioned in Reviewer vuQy's rebuttal were used for computing all 50 local bases for analysis. Additionally, we wish to mention again that the reported 11 seconds include the computation time for the 1st local basis. Because SDEdit is the oldest basic, do-nothing approach, our method has comparable computational complexity except for SDEdit.) Would you reconsider these points?

Authorsrebuttal2023-08-16

We sincerely appreciate the diligent reviewers and AC for their efforts. We would like to kindly ask for missing responses for our initial rebuttal: cKEc, vuQy, and iEgM. In addition, we notice that some reviews are considering our paper as just another diffusion-based editing method neglecting important aspects of our method. We would like to highlight our mathematical and contextual importance: 1. We provide *geometric interpretation of the latent space* and feature space (while the intricate interplay between semantics and the geometric structure within the latent space of diffusion models remains unexplored). 2. We introduce the first method that enables *unsupervised* image editing in both conditional and unconditional models within diffusion models (while contemporary works consider only unconditional models). 3. We show the first instance where semantic editing is possible with just *a single edit at a specific timestep* (while all other editing methods in diffusion model need multiple timesteps). 4. We propose the first method of *directly editing the latent space $x_t$* of the diffusion models. 5. Last but not least, we experimentally demonstrate the *characteristics of the diffusion model by analyzing its basis*. We believe that our paper will significantly enrich future research.

Reviewer cKEc2023-08-16

discussion

Dear Authors, thanks for taking the time to answer my doubts. I'm satisfied with the rebuttal and with the changes you promised to make in the revised manuscript.

Authorsrebuttal2023-08-19

We sincerely thank you for your first review once again. As the discussion phase ends soon, we remain enthusiastic about receiving additional feedback from you. We are ready to accommodate your needs if you find our revised response requires additional clarifications and suggestions. Thank you.

Area Chair CbYX2023-08-21

Thank you for your rebuttal

Dear authors, Thank you for your rebuttal and clarifications -- this supports us as we assess the paper and its reviews during the discussion and decision phase. Thanks, Your AC

Reviewer vuQy2023-08-21

Thank you for the response to the questions. I checked all reviews and responses. Since my questions are addressed, I upgrade my score. However, I agree with the other reviewers that the novelty of the paper should be improved considering its limitations in theory and practical application, i.e. footprints, for a strong acceptance.

Authorsrebuttal2023-08-22

Thank you for dedicating time to share your thoughts. We would like to highlight that our method achieves a comparable complexity, even though efficiency was not our primary focus. (We would like to kindly clarify that the 100 seconds mentioned in the response were used for computing all 50 local bases for parallel transport, not for the default editing process. Additionally, we wish to mention again that the reported 11 seconds include the computation time for the 1st local basis. Because SDEdit is the oldest basic, do-nothing approach, our method has comparable computational complexity except for SDEdit.) Would you reconsider these points?

Authorsrebuttal2023-08-19

We sincerely thank you for your first review once again. As the discussion phase ends soon, we remain enthusiastic about receiving additional feedback from you. We are ready to accommodate your needs if you find our revised response requires additional clarifications and suggestions. Thank you.

Area Chair CbYX2023-08-21

Thank you for your rebuttal

Dear authors, Thank you for your rebuttal and clarifications -- this supports us as we assess the paper and its reviews during the discussion and decision phase. Thanks, Your AC

Reviewer iEgM2023-08-21

I thank the reviewer for their time and effort in making the rebuttal and answering my questions. Overall, I am satisfied with the additional comparison and additional information about time complexity. After having read the other reviews and the author's response, I lean towards acceptance.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC