SEGA: Instructing Text-to-Image Models using Semantic Guidance

Text-to-image diffusion models have recently received a lot of interest for their astonishing ability to produce high-fidelity images from text only. However, achieving one-shot generation that aligns with the user's intent is nearly impossible, yet small changes to the input prompt often result in very different images. This leaves the user with little semantic control. To put the user in control, we show how to interact with the diffusion process to flexibly steer it along semantic directions. This semantic guidance (SEGA) generalizes to any generative architecture using classifier-free guidance. More importantly, it allows for subtle and extensive edits, changes in composition and style, as well as optimizing the overall artistic conception. We demonstrate SEGA's effectiveness on both latent and pixel-based diffusion models such as Stable Diffusion, Paella, and DeepFloyd-IF using a variety of tasks, thus providing strong evidence for its versatility, flexibility, and improvements over existing methods.

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

Peer review

Reviewer Lm7m3/10 · confidence 5/52023-06-10

Summary

The paper introduces Semantic Guidance, a method that enables user control and interaction with text-to-image diffusion models to generate images that align with their intended semantics. The authors highlight the challenge that small changes in the input prompt can result in vastly different images. To address this issue, SEGA leverages the previously proposed composable diffusion approach to allow for fine-grained control over the diffusion process. The experiments showcase the effectiveness of SEGA mostly on human faces.

Strengths

1. This paper studies an important problem, fine-grained control for diffusion models. 2. Experiments on human faces are extensive. Both automatic evaluation and human evaluation are conducted. 3. Results show that SEGA enables subtle edits on human faces, which disentangles the intended edits from factors that should remain untouched.

Weaknesses

1. **Similarity to previous work:** The method is almost identical to Composable Diffusion [1], where different concepts are composed to manipulate the generated image. 2. **Limited scope of evaluation:** Most experiments are done on human faces. It is unclear if the method can be generalized to other object categories, e.g., the height of a building, the width of a bench, or moving an object left/right. 3. **Potential biases:** Several semantic edits target controversial concepts such as gender, hate, and violence. It's better to discuss what is the scope of evaluation here, e.g., what genders are considered and what remains unresolved. [1] Compositional Visual Generation with Composable Diffusion Models. 2022

Questions

1. What are the typical failure cases of the proposed method?

Rating

3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

2 fair

Presentation

3 good

Contribution

2 fair

Limitations

The limitation of the method and evaluation setting needs to be discussed.

Reviewer ZrN76/10 · confidence 5/52023-06-16

Summary

This paper proposes a method for attribute editing of the images generated with Text-to-Image Diffusion models (T2I-DM). This is done in a manner similar to CFG, where extra edit prompts are passed alongside the text prompt to generate the image. Noise predicted from the edit prompt in the denoising process guides the editing of the generated image. The properties of SEGA are extensively reported with qualitative results and on multiple T2I-DMs.

Strengths

1. The key idea is simple, well-motivated, and theoretically backed. 2. The four semantic properties of the method are extensively studied and experimentally backed. 3. Human evaluation results prove the effectiveness of the method.

Weaknesses

1. In Table 1, I suppose the 250 images could or could not contain an attribute (for eg. smile). In that case, the results have been reported by adding smile to 146 images which did not contain it and removing smile from the 93 images that previously had a smiling face. If so, the same protocol could be followed for all other attributes. However, the results for negative attributes are reported only for 3/9 attributes (excluding gender). 2. It would complete the paper if a user preference study against competitive methods such as Wu et al. [1] are also reported. [1] Wu et al. Uncovering the Disentanglement Capability in Text-to-Image Diffusion Models. https://arxiv.org/abs/2212.08698

Questions

Insights about the hyperparameter choices in the supplementary could be presented better. The conclusion statements in each case are missing. What aspect of the generation (image quality, details etc) are controlled by each of the hyperparameter? Perhaps a summary of the hyperparameter analysis could be made part of the main paper, instead of Fig. 6 and its analysis.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

Limitations have been addressed adequately.

Reviewer q8hy7/10 · confidence 4/52023-07-06

Summary

The paper proses a novel approach called Semantic Guidance (SEGA) to edit an image using text only and without fine-tuning the T2I model. It does this by manipulating the noise estimate in the T2I diffusion model. The space of sparse noise-estimate vectors is referred to as the semantic space. The paper justifies various desirable properties of SEGA such as altering fine-grained output details, architecture agnostic (as long as it uses classifier-free guidance), works on both pixel and latent space, no T2I fine-tuning, supports editing of multiple attributes (addition or deletion). Guidance vector is unique per concept and thus can be pre-computed. The value of semantic guidance vector controls the magnitude of semantic concept (e.g. degree of smile). Different concepts/modifiers typically affect different regions of the noise estimate. Experiments in the paper has good qualitative results as well as positive human evaluation on 2 data sets.

Strengths

S1) Clarity S2) Simplicity, elegance, effectiveness, generality of the proposed approach S3) Several useful insights regarding the semantic vector S4) Good qualitative results and human evaluation

Weaknesses

W1) Quantitative evaluation on a large data set. E.g. face attribute detection could be used to do large scale evaluation W2) Not sure how effective SEGA is for more complex manipulations involving multiple objects in the same image e.g. controlling the color of 2 objects independently using referring expressions.

Questions

Please see list of weakness

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

Yes

Reviewer i1VN6/10 · confidence 3/52023-07-06

Summary

This work proposes SEGA (semantic guidance) for text-to-image diffusion models. It is a architecture-agnostic technique and is only applied during the sampling process. In addition to classifier-free guidance, a semantic guidance is applied that pushes noise prediction along the direction of a specific semantic concept.

Strengths

The technique is simple and intuitive. It achieves fine-grained semantic control of image generation via noise-estimate manipulation, therefore an easier process than image-editing models. The results look interesting.

Weaknesses

The human study only evaluates on whether a label is present/absent. No study is done on the quality of images. For example, adding 'glasses' on an image hopefully will not alter the other attributes of the image, i.e. the person still looks the same, the facial expression remains the same, etc. And in addition, the image quality also should not drop significantly. These can be done via user study, or image quality can be measured by inception scores or FID. (I think at least some measure of image quality comparison with the unaltered images is needed.)

Questions

Choices of $c_e$ presented in the paper are all single-word, precise semantic concepts. You tried combining multiple concepts via Eq.10. Have you thought about making $c_e$ a composite concept?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

4 excellent

Limitations

The authors have discussed the limitations and broader impact.

Reviewer 2nZZ5/10 · confidence 5/52023-07-11

Summary

This paper analyzes classifier guidance in a more detailed way. The main idea is to tie the direction of guidance and the semantic, so that multiple guidance can be used. This idea is simple and effective. The paper also presents many small tricks to improve the performance. SEGA is evaluated on the face manipulation task and some qualitative benchmarks with a high successful rate.

Strengths

1. Classifier-free guidance is a very interesting topic in diffusion models, which is the key to improve the performance but neglected by many theoretical papers. 2. The paper tests their method on SD, Paella and DeepFloyd-IF, exhibiting its generality.

Weaknesses

1. The paper presents a solid analysis, but not deep enough for classifier-free guidance. In my opinion, two things are very interesting: (1) What is the range of validity for the linearity of semantics? In word2vec, although the story of linearity is very appealing, but it is not true for many words. In diffusion model, there should be also a range to keep the linearity, which is very important for extending classifier-free guidance. (2) Figure 2 shows only the top 1-5% values are enough to guide the generation. This is an important finding, but it seems like a total statistic from multiple images. In a specific images, where are the top values distributed spatially? What is the relation between these values and initial noise? 2. $\mu$ is used to select the dims with the largest absolute values, which is the largest different between SEGA and usual CFG. However, more discussions about the motivation and quantitative results are needed. The importance of other tricks, e.g. warmup and momentum, are also unknown. 3. The experiments are mainly limited to face-related and qualitative cases. I still think it could be better to compare SEGA with other baselines, e.g. prompt2prompt.

Questions

See weakness.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

See weakness.

Reviewer q8hy2023-08-17

Thanks to the authors for addressing my concerns.

Reviewer i1VN2023-08-20

Response

Thanks to the authors for the rebuttal and the added evaluation and information. It is interesting (and counter-intuitive to me) that SEGA editing improves FID significantly. The observation that "the additional guidance signal for a dedicated portion of the faces frequently removed uncanny artifacts" is an interesting one, and it would be great if this can be reasoned for as well. Overall I maintain my positive rating of weak accept.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC