IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

The advent of Vision Language Models (VLM) has allowed researchers to investigate the visual understanding of a neural network using natural language. Beyond object classification and detection, VLMs are capable of visual comprehension and common-sense reasoning. This naturally led to the question: How do VLMs respond when the image itself is inherently unreasonable? To this end, we present IllusionVQA: a diverse dataset of challenging optical illusions and hard-to-interpret scenes to test the capability of VLMs in two distinct multiple-choice VQA tasks - comprehension and soft localization. GPT4V, the best performing VLM, achieves 62.99% accuracy (4-shot) on the comprehension task and 49.7% on the localization task (4-shot and Chain-of-Thought). Human evaluation reveals that humans achieve 91.03% and 100% accuracy in comprehension and localization. We discover that In-Context Learning (ICL) and Chain-of-Thought reasoning substantially degrade the performance of Gemini-Pro in the localization task. Tangentially, we discover a potential weakness in the ICL capabilities of VLMs: they fail to locate optical illusions even when the correct answer is in the context window as a few-shot example.

Paper

Similar papers

Reviewer 7WKq6/10 · confidence 4/52024-04-26

Summary

This paper introduce IllusionVQA, a novel dataset designed to rigorously test the ability of VLMs to locate and comprehend challenging optical illusions. The authors comprehensively test a wide range of open-source and closed-source VLMs and find some interesting results.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

The dataset is interesting and challenging enough. The results are good.

Reasons to reject

The paper is very similar to hallusionbench. The only difference is that Hallusionbench is binary answer while IllusionVQA is multiple choice questions.

Questions to authors

Please list more difference with Hallusionbench.

Reviewer L4KJ5/10 · confidence 4/52024-05-06

Summary

This paper introduces a new benchmark for evaluating the performance of vision-language models under visual illusions through standard multiple-choice VQA tasks. The authors measure a wide range of model performances under different settings, such as the inclusion of In-Context examples and Chain-of-Thought prompting.

Rating

5

Confidence

4

Ethics flag

1

Reasons to accept

1. The mined dataset of 374 examples can be a useful resource for the community. 2. The study on the effects of in-context examples and chain-of-thought prompting reveals interesting phenomena, such as in-context examples not always being helpful.

Reasons to reject

1. Data contamination and bias in the dataset: GPT-4V is used as a filter in the dataset creation process, which likely results in a dataset biased towards GPT-4V's capabilities. Additionally, the majority of the data is sourced from the web, which likely comes with data contamination issues and is already part of many models' pre-training data. In comparison, prior works like GVIL/Hallusionbench use synthetic data to mitigate this problem. 2. The authors assume that the model should not experience illusions, which is still a topic open for debate. As discussed in GVIL, it remains an open question whether we want models to "experience the same kinds of illusions as humans do, or to faithfully represent reality?" This paper directly uses "the most likely misinterpretation of an optical illusion" (the one humans perceive?) as a negative example, presupposing a specific stance on this issue.

Reviewer 9wgj6/10 · confidence 4/52024-05-10

Summary

This paper presents a novel dataset aimed at evaluating Vision Language Models (VLMs) on their ability to understand and localize optical illusions, featuring two principal tasks: comprehension and soft localization. These tasks test VLMs' proficiency in interpreting complex visual phenomena that are typically misleading to human vision. The dataset has been rigorously tested on various state-of-the-art VLMs, including GPT4V and Gemini-Pro. While these models generally perform well in standard object recognition, they exhibit significant difficulties with optical illusions, especially in tasks demanding intricate visual and spatial reasoning. Overall, this paper is well-structured and clear, with methodical descriptions of experimental methods, dataset construction, and results, enhancing its accessibility and understanding. However, this paper lacks detailed descriptions of data sources, raising concerns about its diversity and applicability for practical VLM applications.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

The IllusionVQA dataset introduces a unique challenge to the field of VLMs by focusing on optical illusions. This dataset differs from traditional VQA datasets by specifically testing the models' comprehension and spatial reasoning abilities through its two main tasks: comprehension and soft localization. These tasks assess VLMs' capability to interpret and localize elements within images that are inherently deceptive, extending beyond basic object recognition to probe deeper into advanced visual processing. Initial testing on the IllusionVQA dataset highlights significant shortcomings of contemporary VLMs, including GPT-4V. The underperformance compared to near-perfect human results underscores the current gap in VLMs' ability to handle complex visual inputs.These findings point to a pressing need for advancements in VLM architectures and training methods.

Reasons to reject

This paper lacks clear descriptions of the data sources for its optical illusion images, raising concerns about the diversity and representativeness of the dataset. There is no detailed methodology on ensuring data comprehensiveness, which is crucial for a varied visual category like optical illusions. Furthermore, the study does not discuss whether the illusions are representative of real-world scenarios, which would be vital for practical applications of VLMs. Addressing these issues could enhance the dataset’s utility and relevance, providing a stronger benchmark for VLMs in practical environments.

Reviewer XYvp7/10 · confidence 4/52024-05-24

Summary

This paper introduces IllusionVQA, a dataset for testing the ability of VLMs to locate and comprehend challenging optical illusions. The dataset consists of two tasks IllusionVQA-Comprehension and IllusionVQA-Soft-Localization. It is an interesting and novel problem and the community will be benefited from this perspective.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

1. The dataset includes a diverse set of 12 distinct categories of optical illusions. 2. The proposed problem is relatively hard for current models. The authors evaluate a wide range of open-source and closed-source VLMs. Experiments show that VLMs can locate ordinary objects accurately but struggle with optical illusions. 3. The paper introduces a novel "soft localization" task that tests VLMs' ability to differentiate geometrically impossible objects from ordinary objects, which has not been explored in previous studies.

Reasons to reject

1. Potential to scale up the dataset is limited. The dataset size is relatively small, but it is totally understoodable, due to the difficulty to find additional high-quality optical illusions that met the inclusion criteria, which limits the scale of the dataset. Although the authors mentions synthetic optical illusion generation could be a potential way, but current image generation models have limited capabilities to follow such instructions.

Reviewer XYvp2024-06-01

Thank you for addressing the comments and questions.

Reviewer L4KJ2024-06-06

Thank you for the response! However, I still think GPT4V bias is a concerning issue for a benchmark-centric paper and am therefore keeping my scores.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC