What makes a good metric? Evaluating automatic metrics for text-to-image consistency

Language models are increasingly being incorporated as components in larger AI systems for various purposes, from prompt optimization to automatic evaluation. In this work, we analyze the construct validity of four recent, commonly used methods for measuring text-to-image consistency - CLIPScore, TIFA, VPEval, and DSG - which rely on language models and/or VQA models as components. We define construct validity for text-image consistency metrics as a set of desiderata that text-image consistency metrics should have, and find that no tested metric satisfies all of them. We find that metrics lack sufficient sensitivity to language and visual properties. Next, we find that TIFA, VPEval and DSG contribute novel information above and beyond CLIPScore, but also that they correlate highly with each other. We also ablate different aspects of the text-image consistency metrics and find that not all model components are strictly necessary, also a symptom of insufficient sensitivity to visual information. Finally, we show that all three VQA-based metrics likely rely on familiar text shortcuts (such as yes-bias in QA) that call their aptitude as quantitative evaluations of model performance into question.

Paper

Similar papers

Reviewer Gt1f7/10 · confidence 3/52024-04-12

Summary

Research question of this paper is: "What visio-linguistic properties make a good text-to-image consistency evaluation metric?" The authors identify linguistic metrics (readability, complexity and length) and visual metrics (imageability, concreteness, overlap with large scale image benchmarks), and then meta-evaluate four evaluation metrics: TIFA, VPEval, DSG and CLIPScore. Finally, the authors analyze the pairwise correlation across metrics and find that three metrics provide new information to CLIPScore.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

This paper can be a good contribution to COLM. It is well written, well presented and provides an important analysis: How well do we progress in evaluating text-to-image matching? I especially find the identified visio-linguistic properties interesting and necessary (even though incomplete). I especially enjoy reading Ablation: Filling in the gaps section, since it provides insights on how to design better text-to-image evaluation measures (e.g. incorporating information from both modality).

Reasons to reject

- No major weaknesses are identified. My only minor concern is the abstractness of some visual metrics, such as imageability. - Do you mind comparing your work against the concurrent preprint: "Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with (TS2) " (https://arxiv.org/pdf/2404.04251v1.pdf). - You should settle the real title of your work: Is it: 1) What makes a good metric? Meta-evaluating automatic metrics for text-to-image consistency 2) What makes a good metric? Evaluating automatic metrics for text-to-image consistency

Questions to authors

see above

Reviewer KjKK5/10 · confidence 4/52024-04-29

Summary

This paper presents a study to explore the four existing automatic evaluation metrics (CLIPScore, TIFA, VPEval, and DSG) for text-to-image generation. It first designs a set of desiderata, and found none of them meet all the desiderata. And also deliver several findings for the existing metrics: relying much on the text, VQA-based metrics have very skewed question distributions.

Rating

5

Confidence

4

Ethics flag

1

Reasons to accept

1. It defines a set of classes for desiderata for metrics, which might be useful for the future metrics proposals. 2. It also proposes a couple of data augmentation ways to fill the gaps. 3. And it delivers some interesting findings.

Reasons to reject

1. My impression to this work is - it empirical explores the four existing metrics in the text-to-image evaluation, and finds some common issues for these metrics, while it does not give a solution yet. It is an interesting paper to explore the existing metrics in the text-to-image evaluation. And the findings in the paper is not surprised. As CLIPScore and other three VQA-based metrics are all pretrained models, they should inherit the bias from their pretraining. For the second part of this work, it might be better to deliver a new metric or evaluation method based on the proposed desiderata, which may look a more complete work. Overall, my evaluation to this work is a borderline work, can be either 5 or 6.

Questions to authors

See my rejection part.

Reviewer V1Vf6/10 · confidence 4/52024-05-10

Summary

This paper evaluates a collection of recent LM + VQA based text-to-image evaluation metrics, and examined whether those approaches are faithfully evaluating some basic desiderata of an ideal text-to-image evaluation metric. Particularly, the author evaluates whether those metrics are human interpretable, and their sensitivity to text properties and image properties, and whether they were robust to known shortcuts. The analysis of those evaluation metric has suggested that those LM + VQA based metric is providing complementary information about the evaluation to the more popular CLIP Score, but they highly correlates between themselves.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

1. The paper is overall clear and the research ideas are straight-forward. The research finding is kind of interesting, suggesting that three existing VQA based evaluations for text-to-image generation are highly correlated, and all having the similar yes-bias. 2. The analysis done for each comparison is convincing, with strong supports from the experiments. 3. The design of experiments are reasonable and solid in most cases, and the ablations are comprehensive.

Reasons to reject

1. While it has been a pleasant time reading the paper, it does not carry very surprising or significant research discover that reveals any new research direction. It is completely okay to have a small and focused research paper like this but I wouldn't consider it to be the top papers for CoLM. 2. To improve the reading experience, maybe It would be good to have a figure illustrating a generation example and show how it got evaluated by each metrics. It could be even better if that example can also tell the story of the flaws identified with the VQA based approach 3. The proposal of the desiderata future text-to-image metrics should consider is a bit abstract. Especially the "iii) nice-to-haves" is very vague and not well defined.

Questions to authors

1. Section 3.4 paragraph 1 "# of generates questions correlates strongly to the text-image consistency metrics" is confusing to me and I don't see how it leads to the conjecture that VQA component can be omitted from the evaluation pipeline. Let's say now we remove the VQA component, and only have the question generation part, how is model going to compute the metric score for text-to-image consistency? Are you suggesting that we directly measuring the "# of generated question" as consistency score? There is no way that this should be the future design of text-to-image eval metric. 2. Most analysis in this paper are statistically, could you also provide some intuition and concrete illustrating examples? For instance, when experiments are designed to show that VQAbased scores are not correlated strongly with CLIPScore, what would be the possible justification over this observation? Could you show examples to justify the reason?

Reviewer Gt1f2024-06-02

My positivity remains

Dear authors Many thanks for providing an answer, and especially a comparison against the very similar concurrent work. I still believe your paper should be of interest to the community. It works on a relatively under-explored field: evaluating the evaluation metrics. Everyone complains that the evaluation metrics are imperfect but not many attempts to change it. There, I keep my original score.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC