Thank you for your response
Regarding “dataset”, “metric” and “memorization” sections being part of a coherent paper, I respectfully disagree with the authors: the “metric investigation” section, uses the dataset to study several metrics, but rather than focusing on why existing metrics do not correlate with Human Error and if that reveals a limitation in the Human Error metric itself or the existing metrics (since these metrics are backed by their own user studies and theories), it quickly takes Human Error as ground truth and moves on to improving metrics (hence a separate paper in my opinion); b) the “memorization” section does not use the proposed dataset or the Human Error metric, it does not even mention them.
1.1. You claim “each participant sees a total of 250 images from ImageNet (which has 1000 distinct classes), making it essentially impossible to learn any diversity within any class, and thus forcing a focus on fidelity”, the issue is not that your participants will learn to focus on diversity or not, it is that they might already associate rarity with fakeness, so you need to provide evidence for the claim that “... forcing a focus on fidelity”: one way to do so is to report Human Error rate on rare real samples versus common real samples (e.g. using Rarity Score, Han 2022), and see if the error rate is the same on both sets. Lack of such experiments, and any discussion of this potential issue in the paper, is why I do not think you can claim that “our methodology accounts for and negates such an effect” at present.
1.2. The parrot/crow is of course just an extreme example to clarify my point, within each ImageNet class you can find unlikely as well as likely samples.
1.3. You claim “if participants were confusing fakeness with unlikeliness, this effect would be consistent across models and therefore would not alter rankings”, I don’t see why the effect would not alter rankings: if participants confuse fakeness with rarity, they will make less mistakes on models that generate more rare samples (because they will just flag them as fake based on the samples being unlikely), so your score would unfairly rank models that are more diverse worse.
1.4. I don’t think being similar to HYPE or aided by psychophysicist answers any of my very specific concerns. What is lacking here is quite straightforward, as I explained in 1.1, you should provide a control experiment, otherwise the dataset – which I want to emphasize that I think is valuable and interesting – will incentivise a series of misleading and incorrect followup works.
2.. I don’t understand what this means: “The Normalized Error Rate accounts for the varying difficulty of tasks over different combinations of (dataset, model).”, please elaborate.
3.1. You claim “Figure 2 clearly shows that diffusion models score the highest human error rates, and that GANs often score lower human error rates yet achieve a better FID ranking”, but in Figure 2 CIFAR10 the diffusion model is ranking best in terms of FID too, so FID is not unfair. Same in ImageNet. My point is by just looking at Figure 2, there is no concrete evidence to back “unfairness”. You need to be more specific about what is the mathematical definition of “unfair” in your work, and how you measure it. For example, you could define unfair as low correlation between FID and Human Error Rate, and then report the correlation coefficients and claim that the correlation is higher for GANs, but lower in Diffusions.
3.2. You claim “Table 1 summarizes the statistical significance”, yet you report no significance test results, it is unclear what the ordering in this table are based on. Table 11 also does not clearly show “unfairness” towards diffusion models.
6.. I understand your motivation, but what I still do not understand is what “fundamental” means in this context. In any way, I do not consider this a main concern, I acknowledge that whether something is a “fundamental” flaw, or simply a lack of correlation between some metrics, is subjective. I appreciate the additional explanations by the authors.
7.. If computational restrictions do not allow you to sufficiently support the claim that DINOv2 is better than Inception, please avoid making that claim in your abstract. If Figure 6 is able to sufficiently support this claim, please elaborate.
8.. I agree that there might be interesting connections, but lack of a study on those connections makes the memorization results inconclusive.
Han, Jiyeon, et al. "Rarity score: A new metric to evaluate the uncommonness of synthesized images." ICLR 2022.