ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward -- the first general-purpose text-to-image human preference reward model -- to effectively encode human preferences. Its training is based on our systematic annotation pipeline including rating and ranking, which collects 137k expert comparisons to date. In human evaluation, ImageReward outperforms existing scoring models and metrics, making it a promising automatic metric for evaluating text-to-image synthesis. On top of it, we propose Reward Feedback Learning (ReFL), a direct tuning algorithm to optimize diffusion models against a scorer. Both automatic and human evaluation support ReFL's advantages over compared methods. All code and datasets are provided at \url{https://github.com/THUDM/ImageReward}.

Paper

Similar papers

Peer review

Reviewer 6uei6/10 · confidence 4/52023-07-06

Summary

The paper presents a human preference model for text-to-image generation and a method to enhance text-to-image models using this preference model. To achieve this, the authors develop a human preference annotation pipeline and create a dataset consisting of generated images and human ratings. The proposed model is trained to predict human preference rankings, and experimental results indicate that it aligns better with human preferences than existing automatic measures. Additionally, the paper introduces a learning method to fine-tune a diffusion model using the human preference model. The experimental result demonstrates that the text-to-image generation model, when adjusted with the proposed method, is preferred by human annotators.

Strengths

- The paper's clear contribution is the human preference dataset, which features high-quality annotations from a professional data annotation company. - The experiment suggests that ImageReward outperforms popular measures like FID and CLIP scores in evaluating text-to-image generation. - By training the text-to-image model using the human preference model, the authors achieve improved performance in both automatic and human evaluations.

Weaknesses

- The paper exceeds the page limit. The authors should carefully follow formatting instructions and revise the manuscript accordingly. - The description of ImageReward training lacks detail, which may make reproduction difficult. - Information regarding the human evaluation of experiments in Section 4.2 is missing. - Certain aspects are unclear. For example, the intention of Figure 7 is unclear for me. What image generation problems are demonstrated in the examples? The numbers in Table 2 are unexplained. What agreement measure is used?

Questions

- Will the entire human preference dataset be made publicly available? The repository looks to provide only test set. - Can the authors clarify the unclear points mentioned in the Weaknesses section?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

The limitations section addresses three main issues: - Annotation scale, diversity, and quality - The heuristic nature of RM training - The lack of theoretical background for training the diffusion model with RM feedback The authors provide meaningful suggestions for future research directions.

Reviewer 5G858/10 · confidence 5/52023-07-06

Summary

This study presents ImageReward, a general-purpose text-to-image human preference reward model. They have collected 137k expert preference dataset, which contain a lot of rating and ranking annotation. Additionally, the authors propose Reward Feedback Learning (ReFL), a direct tuning algorithm designed to optimize diffusion models. The performance of ImageReward surpasses that of existing models and metrics.

Strengths

1. The paper is very well presented with clear paper writing and good demonstration. 2. The idea is very novel, which explores directly using the human preference as the supervision signals to tune the pretrained text-to-image generation models 3. they also collect a large-scale high-quality human preference dataset, which can inspire many future works.

Weaknesses

BLIP, being outdated, falls short in generating accurate and comprehensive image descriptions. An alternative approach is to employ more recent models like MiniGPT-4 or LLaVa, which have the potential to produce superior reward scores. These advanced vision-language models not only comprehend the objects within the image but also grasp the emotional and artistic aspects. It would be beneficial if the authors could include a comparative analysis involving these models to further strengthen their findings.

Questions

you may also incorporate evaluation metrics and more complex prompts (e.g. with more spatial grounding") to collect the human feedback.

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

4 excellent

Contribution

4 excellent

Limitations

it is well discussed in the paper

Reviewer JPbK6/10 · confidence 5/52023-07-07

Summary

The paper explores human preferences and introduces an ImageReward mechanism, which can be employed for evaluating text-to-image generation. The study further enhances the performance of existing text-to-image generation models through a Reward Feedback Learning (ReFL) approach. The experiments primarily focus on the alignment between the proposed ImageReward and human preference/judgment, which is annotated by real individuals.

Strengths

The problem addressed in this paper is vital as it aims to judge the alignment between text-to-image generation models and human preferences during both training and evaluation. The proposed ImageReward model could effectively guide both the training and evaluation processes of text-to-image generation models.

Weaknesses

There are some concerns about experiments including the considered generative models and the results of the correlations with human preferences/judgments. Please refer to more details in the following.

Questions

1. The number of generative models considered in the experiments seems insufficient. With only six text-to-image generation models used for the evaluation of the proposed ImageReward, the ranking results, as demonstrated in Table 1, may be easily consistent. Moreover, even a slight variation in ranking can lead to significantly different correlation values. To enhance the persuasiveness of the experiments, it would be better to include a more diverse range of generative models, which is more convincing. 2. In Table 1, the results show that the Spearman correlation between the zero-shot FID score and human judgment is only 0.09. With this relatively low correlation, does it suggest that there is little relevance between the FID score and human preferences?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors discuss the limitations and broader impact in the last two sections of the paper and propose potential solutions and outline future directions for further research and development.

Reviewer 7vWQ6/10 · confidence 3/52023-07-08

Summary

The paper introduces ImageReward, a new dataset of human preference over generated images given a text prompt. Human preference for images is rated across three dimensionalities: text alignment, image fidelity, and harmlessness. Using the dataset, they train a reward model to score the generated image and text prompt pair. The reward model consists of a small trainable MLP head over BLIP text-image features. The paper further proposes a baseline method of fine-tuning the generative diffusion model using the score model to increase the generated image alignment with human preferences. The fine-tuning loss is a weighted sum of standard diffusion loss and the one from the scoring model (only at the predicted images at lower timesteps of diffusion).

Strengths

The annotated dataset of human preference is one of the first datasets of this kind and scale. This will help in further research in both text-to-image model evaluation and improving the generations with higher human preference. Both the reward scoring model and fine-tuning method based on the score model are shown to work on par of better than existing baselines. The paper consists of extensive analysis and details regarding the dataset annotation, scoring model, and its comparison to recent methods.

Weaknesses

1. One limitation of the dataset might be that it becomes less relevant as the generative models improve. Given that the dataset only consists of generated images with their corresponding text prompt and not a plausible ground truth image for the prompt with the best possible human preference score. 2. Only fine-tuning the model on 0-10 timesteps with the scoring model doesn't seem optimal. Specifically, in cases of object omission, the layout has already been decided in the initial stages of diffusion; thus, the reward score guidance at later stages might not be effective. Is there any ablation or analysis regarding what metrics among fidelity and text alignment improve the most? 3. In Eq2, phi is implemented as a ReLU function, as mentioned in line 260. Probably this should be ReLU over the negative of the score function. Because the higher the score, the better. Or is my understanding incorrect? 4. It would be great to expand on the evaluation setup and metrics, which are sometimes not very clear. (a) In line 122, are the 100 real user test prompts different than the ImageReward dataset used to train the scoring model? Similarly, 466 and 371 prompts in Table 3. (b) How is the "filter" evaluation metric calculated in Table 3? Does this calculate the number of times the model didn't select the worst image in top-k? (c) In Table 4, it's unclear how the evaluation numbers are reported. Does it denote the #winrate for each method out of total N samples (N being the sum of the column), or is it a binary comparison of each method vs the baseline? How many generated images per prompt over the 466/77 prompts were used for the evaluation? (d) Some of the baseline, e.g., reward weighted fine-tuning method, performs worse than the baseline in Table 4. Is there any analysis regarding that?

Questions

Minor point: some grammatical issue in line 290, 302 and 360

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

yes

Reviewer kP497/10 · confidence 4/52023-07-09

Summary

This paper aims to improve text-to-image (T2I) from human preference feedback. They first collect a human rating dataset to train their human preference model, ImageReward. With the reward model, they further optimize a pre-trained T2I via the proposed Reward Feedback Learning (ReFL). The experimental results indicate that their ImageReward is more robust than the widely-used CLIP-Score, and ReFL-optimized T2I also performs better than baselines.

Strengths

+ This paper is well-written and easy to follow. + The collected dataset is valuable for the V+L community as well as the trained ImageReward, which can help various visual generation tasks (not only T2I). + The proposed ReFL pipeline can keep improving the T2I model. Both automatic metrics and human evaluation support the superior performance of their framework. + They provide lots of qualitative examples and detailed discussion in the supplementary.

Weaknesses

+ The novelty can be an issue since the human preference-trained reward model and feedback learning are already introduced in large language modeling (LLM). It looks like they just apply the same pipeline from LLM to T2I. + It is not easy to collect large-scale human preferences for reward model training. Is there a more efficient way to build ImageReward instead of fully relying on human annotations? + A detailed analysis of ImageReward should be considered. For example, how many human preference pairs can lead to how well ImageReward and then how well the final optimized T2I model is.

Questions

Please see the Weakness

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

4 excellent

Limitations

Since their pipeline relies on the collected labels, the human annotation can contain ethnic issues, which they also discuss in Appendix B.

Reviewer 6uei2023-08-12

Thanks to your response

Thank you for your responses. The responses addressed my concerns. I also appreciate additional dataset contribution. I am increasing my score.

Reviewer 7vWQ2023-08-18

Thanks for the response.

Thanks for the detailed response in the rebuttal. It addresses my concerns. I am keeping my score of weak accept.

Area Chair 4gNY2023-08-19

Dear authors, Thank you for your taking the time to respond to the comments. Dear Reviewer kP49, After reading the authors' response, do you have any additional thoughts? Best, AC

Area Chair 4gNY2023-08-19

Dear authors, Thank you for your taking the time to respond to the comments. Dear Reviewer JPbK, After reading the authors' response, do you have any additional thoughts? Best, AC

Reviewer JPbK2023-08-19

Thanks for the response. It addresses all of my concerns. I will raise the score.

Area Chair 4gNY2023-08-19

Dear authors, Thank you for your taking the time to respond to the comments. Dear Reviewer 5G85, After reading the authors' response, do you have any additional thoughts? Best, AC

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC