Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)

With advances in the quality of text-to-image (T2I) models has come interest in benchmarking their prompt faithfulness -- the semantic coherence of generated images to the prompts they were conditioned on. A variety of T2I faithfulness metrics have been proposed, leveraging advances in cross-modal embeddings and vision-language models (VLMs). However, these metrics are not rigorously compared and benchmarked, instead presented with correlation to human Likert scores over a set of easy-to-discriminate images against seemingly weak baselines. We introduce T2IScoreScore, a curated set of semantic error graphs containing a prompt and a set of increasingly erroneous images. These allow us to rigorously judge whether a given prompt faithfulness metric can correctly order images with respect to their objective error count and significantly discriminate between different error nodes, using meta-metric scores derived from established statistical tests. Surprisingly, we find that the state-of-the-art VLM-based metrics (e.g., TIFA, DSG, LLMScore, VIEScore) we tested fail to significantly outperform simple (and supposedly worse) feature-based metrics like CLIPScore, particularly on a hard subset of naturally-occurring T2I model errors. TS2 will enable the development of better T2I prompt faithfulness metrics through more rigorous comparison of their conformity to expected orderings and separations under objective criteria.

Paper

References (58)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Xy9h6/10 · confidence 4/52024-07-07

Summary

The authors propose T2IScoreScore (TS2), a benchmark and set of meta-metrics for evaluating text-to-image (T2I) faithfulness metrics. Compared to existing relevant benchmarks, TS2 has higher image-to-prompt ratios, which allows users to organize semantic error graphs (SEGs), where each edge corresponds to a specific error with respect to the prompt that a child image set possesses but its parent images do not. Based on SEG, the authors evaluate T2I metrics, including embedding-based (CLIPScore/ALIGNScore), QG/A-based (TIFA/DSG), and caption-based (LLMScore/VIEScore) metrics, by how the metrics properly order and separate images. Different metrics show different advantages, and the authors highlight that simple embedding-based metrics outperform other more computational metrics in separation criteria.

Strengths

1. Introduction of meta-metric benchmarks of recent T2I metrics, including collection of large image-text pairs and semantic error graphs. 2. Comprehensive experiments, including using different VLM backbones for TIFA/DSG.

Weaknesses

**1. Simply treating QG/A metrics as score regressors.** One of the major motivations behind using QG/A metrics, even though they are computationally expensive, is that they divide multiple aspects of prompts and provide comprehensive skill-specific performances of the T2I model. While the QG/A metrics can be used as a score regressor, comparing them with embedding-based metrics ignores their biggest advantages. This background and limitation needs to be clarified in introduction section, otherwise this could misleading readers who are new to T2I metrics.

Questions

See weaknesses

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

There is no significant negative societal impact.

Reviewer h3XZ7/10 · confidence 4/52024-07-13

Summary

This paper presents a rigorous evaluation for text-to-image alignment metrics. This is primarily done by introducing a dataset with several images for each prompt, allowing the construction of semantic graphs that can be used to measure the accuracy of the alignment metrics. From the analysis on the benchmark, a major conclusion is that CLIPScore provides an excellent tradeoff (or is at least on the pareto-optimal frontier) between speed and alignment. VQA-based metrics (e.g TIFA, DSG) while improving over CLIPScore in many cases, come with much higher costs (in some cases orders of magnitudes higher), highlighting important considerations for text-image alignment methods.

Strengths

The dataset collected in the paper is quite valuable, and would be useful for evaluating text-image alignment metrics in the future. The methodology in the paper also seems quite sound to me. I also think the analysis in the paper is very sound, highlighting the cost of running the evaluation metric is an important aspect which is often missed in these methods. The paper is also written very clearly, and is easy to read.

Weaknesses

In terms of models/methods evaluated, I see 2 notable omissions: human-preference models such as ImageReward would also be a good addition since they might also capture some notion of text-image alignment while also cheap to use. Another good addition would be VQAScore (much more recent), but seems to show extremely strong results on several image-text matching benchmarks, while not being nearly as expensive as the Question-Generation methods (e.g. TIFA, DSG). Minor: While I totally agree with the issue of methods evaluating on their own proposed test set, the column "ad-hoc" in Tab. 1 makes little sense without an explanation (line 85 seems too limited). Either a more complete explanation should be given, or it would be better to replace this column with something more objective/concrete. [a] Lin et al. "Evaluating Text-to-Visual Generation with Image-to-Text Generation", 2024

Questions

Overall, I really like the paper and would like to see this accepted. I have a couple of questions for my curiosity. While 2.8k prompts is still a lot more than TIFA, DSG, is this size still not too small/limited diversity to be able to comprehensively make conclusions about alignment methods? Is there any idea of what 'human-performance' would look like on these benchmarks, and if existing models are very far off from that?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

No major concern.

Authorsrebuttal2024-08-09

Thank you for clarifying. These are great points: 1. TS2 + and orthogonal quality-only eval might actually tease out the degree to which human annotators attend over each consideration by directly comparing their preference correlations to each metric under different annotation schemes. A really cool idea might be to treat those two considerations as principal components a metric could be interpolated between, or enabling a search for a jointly optimized single "best metric." We will definitely investigate adding ImageReward to the final leaderboard given this. 2. Yes, this is another stimulating direction for future work. Using this human baseline also has implications for direct "VLM-as-a-judge" metrics that ask them to provide Likert scores, as only "superhuman Likert-assigning" VLMs would be sufficient to beat humble CLIPScore. 3. Agreed, we will make sure to clarify that QG/A metrics have this interpretability advantage, and in particular that this consideration, alongside TS2 evaluation, helps users choose a metric based on scenario and needs. Eg., in an interactive app (relatively low analysis throughput rate) cost is less of a consideration, and a human would benefit from interpretable analysis. Whereas for an online reward/feedback model or a supervisory post-filter for image generation, cost is an important consideration and interpretability isn't. Grounding metric selection in all of these considerations is best; TS2's contribution is in capturing one important consideration well. Thanks for the further stimulating discussion!

Reviewer cvSS7/10 · confidence 4/52024-07-14

Summary

The paper proposes an evaluation framework for holistically assessing text-to-image (T2I) evaluation methods. Since most of them are primarily established through simplistic correlational evidence and only compared to the CLIPScore baseline, this approach presents a more detailed way of assessment and also benchmarks existing promising evaluations. While there isn't a single clear winner, results suggest that CLIPScore is still a very competitive candidate and is especially successful when considering the much lower compute costs.

Strengths

- important contribution to investigate and improve automatic T2I evaluation strategies - interesting dataset construction which appears to build a more challenging test bed compared to previous methods, allowing more detailed insights and distinctions - thoughtful discussion of the results & exploration of limitations - very informative figure 2

Weaknesses

- I like the general setup but I'm unsure about the accuracy of the "number of errors" counting system. I'll give two examples from within the paper. Take the last example in Figure 3 with the prompt "A gray elephant and a pink flamingo". An image with two flamingos is categorized as containing one error because there is no elephant. However, if there additionally was an elephant, it would still have an error since there are two instead of one flamingo. So one could argue that there are in fact two errors: a missing elephant and an additional flamingo. Or you say it's one error because there is one animal that is a flamingo but should be an elephant. So this is inherently ambiguous. However, in any ranking solution, this can actually matter quite a lot, so I'm worried that this introduces noise into the analysis process that is hard to reason over. I'm wondering to what extent this is taken into account by the design and how sensitive the results are to this. (Second example to illustrate from Figure 1 SEG: it's noted that when the shirt is not green, it's counted as one error. What if the shirt was additionally also suddenly a hoody. Is that then two errors or still only one? When the boy is gone overall, it's two errors because the shirt isn't green and there is no boy -- but what if there was now a grey shirt in the picture?) - Given that the evaluation framework provides many different evaluations for varying setups and (as discussed in the paper) those might come with their own biases, what is the recommendation to those who are thinking about using this framework for when they can call their metric successful? Defining an overall evaluation aggregate might also help with the adoption of the framework. - (Minor: It's hard to see which numbers are italicized in the results table.)

Questions

None

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Sufficiently addressed.

Reviewer vRNq5/10 · confidence 3/52024-07-21

Summary

The paper introduces T2IScoreScore (TS2), which aims to evaluate how good newly developed text-to-image (T2I) evaluation metrics/methods are. The authors formalize the task of evaluating t2i metrics as their abilities to *order* images correctly within SEGs.

Strengths

1. The authors identify a very important task -- to evaluate T2I metrics. The introduction of T2IScoreScore and the use of semantic error graphs (SEGs) to evaluate T2I faithfulness metrics are novel and innovative. 2. Experiments are good. The methodology is rigorous, and the experiments are well-designed to test the core claims of the paper.

Weaknesses

1. Limited Scope: The evaluation is primarily focused on a specific subset of T2I models (many variants of SD, and Dalle-2) and metrics. Expanding the scope to include a broader range of models and datasets would strengthen the generalizability of the findings. Potentially should consider other synthetic images from models such as OpenMUSE or aMUSEd (https://huggingface.co/blog/amused) with totally different generation architectures than diffusion, etc. Alternatively, text-2-image is an old task, even GAN and VAE can probably have image distribution other hand SD and Dalle-2 which heavily depends on CLIP. 2. Intrinsic Bias: The reliance on rank-correlation metrics, which have intrinsic biases, might affect the evaluation results. A more thorough discussion of these biases and potential alternatives could enhance the robustness of the conclusions.

Questions

Instead of pairwise ranking/comparisons, which might not always be robust, have the authors design multi-images ranking instead of only two images for pairwise comparisons? Something like bradley-terry style ranking could make the evaluation more robust.

Rating

5

Confidence

3

Soundness

2

Presentation

3

Contribution

3

Limitations

n/a

Reviewer h3XZ2024-08-09

Some Comments/Suggestions

I thank the authors for their reply, I have no major concerns left about the paper, and I see that all the other concerns of the reviewers are satisfactorily addressed. That said, I have a few comments that the authors may wish to think about: 1) Human-Preference Models: I agree that human-preference models are not immediately clear about what exactly they evaluate. They also start from CLIP/BLIP models, and then finetune it on data which capture some mixture of visual quality (i.e artifacts), aesthetics, and prompt following all at once. For instance, if a vital object from the prompt is missed in the image, the user would naturally rate it low. Similarly, if the image has a lot of artfiacts but follows the prompt well, it is unlikely to do well on comparisons. Depending on the guidelines, annotation protocol, you can get models that result in very different behaviors. For instance, in the VQAScore paper (disclaimer: I have no connection to it), ImageReward outperforms both CLIP, BLIP on most benchmarks (Tab. 4) and is even doing reasonably well on Winoground, Eqben (Tab.3 ). Therefore, I would not dismiss them as solving an entirely different task, and adding TS2 as an additional eval benchmark for these models would be a good idea. 2) Human-Evaluations: I think the authors make an excellent point that humans doing Likert scoring of image-prompt accuracy might actually not outperform existing metrics. This is a useful pointer (since some prior works recommend Likert scoring of image-prompts as a good strategy to evaluate models[a]) in performing more rigorous user studies for evaluating text-to-image models. 3) QG/A metrics doing more than a single score regression: To reviewer Xy9h, the authors point out that QG/A metrics are proposed claiming superior correlation with human judgement on various benchmarks. While I agree with this, the authors should look at Fig. 1 of TIFA which clearly makes the claims of "fine-grained", "accurate", "interpretable". Of course, the fine-grained/interpretable aspects are the hardest to evaluate and justify, therefore papers will inevitably resort to maximizing performance/correlation on benchmarks to justify the method. That does not mean the other aspects of the method are invalid/absent, they are just insufficiently evaluated (beyond a few qualitative examples). Therefore, I would suggest the authors that they acknowledge the strength of QG/A methods, while providing a fair assessment of their shortcomings (which is already there in the paper). I hope the authors can think about these aspects and make the additions/modifications that they deem fit for the camera ready/benchmark leaderboard. [a]: Otani et al. "Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation", CVPR 2023

Reviewer cvSS2024-08-09

I thank the authors for their response. What you're saying makes sense to me, especially when it comes to the error counting matter. I just want to reiterate on one part of my prior review which is on defining when this framework establishes "success". I understand that this framework provides a detailed holistic overview on a range of interesting dimensions (see Table 2). However, I'm wondering whether there is a recommendation for researchers who want to use this framework to choose the best-performing solution. Are there specific rankings/dimensions that are most diagnostic for overall performance? (And for adoption in the broader community, having an overall score that accumulates the individual results in Table 2 might help with adoption in the community. Do you have a suggestion what this score might be?)

Authorsrebuttal2024-08-09

Thanks for a quick follow up! To give a more committal answer your question: Intuitively, the walk-based spearman correlation metric is probably the best choice for an overall score, as it captures the core desideratum of "able to correctly compare similar images by structural differences." For this desideratum, higher is always better. While the other two scores do matter---it is important that adjacent nodes be statistically significantly separated---it is less clear that higher is always better, vs anything over a threshold being sufficient. Thus a good recommendation might be to rely on the walk ordering score (which is also the main novel contribution here) and to treat the separation scores as a secondary consideration. In other words, par performance on the separation metrics alongside significant gains on ordering would be a very positive development, whereas significant improvement along separation with a loss in ordering could be negative---exact reverse ordering of the nodes with statistically significant node separation would get high delta scores, but be very bad. This is the recommendation and justification we will provide in the camera-ready: **a metric is clearly superior to others when it presents significantly higher ordering (particularly over the hard *nat* subset) without a significant drop in separation scores**---equivalent separation is sufficient. Ultimately, this recommendation is a judgement call and its main grounding is the aforementioned theoretical analysis (high ordering score always captures a good correlation to the scores, whereas high separation score can be present, even when the ordering is reversed). We will use this analysis to justify our recommendation to researchers in the conclusion section of the camera ready.

Reviewer cvSS2024-08-13

I thank the authors for their clarifications in the rebuttal. They sufficiently addressed my concerns and I updated my scores accordingly.

Reviewer Xy9h2024-08-12

I appreciate authors' response and decided to keep my current score. Please make sure to incorporate the points you mentioned in the next version if accepted.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC