Summary
This paper presents a novel dataset aimed at evaluating Vision Language Models (VLMs) on their ability to understand and localize optical illusions, featuring two principal tasks: comprehension and soft localization. These tasks test VLMs' proficiency in interpreting complex visual phenomena that are typically misleading to human vision. The dataset has been rigorously tested on various state-of-the-art VLMs, including GPT4V and Gemini-Pro. While these models generally perform well in standard object recognition, they exhibit significant difficulties with optical illusions, especially in tasks demanding intricate visual and spatial reasoning.
Overall, this paper is well-structured and clear, with methodical descriptions of experimental methods, dataset construction, and results, enhancing its accessibility and understanding. However, this paper lacks detailed descriptions of data sources, raising concerns about its diversity and applicability for practical VLM applications.
Reasons to accept
The IllusionVQA dataset introduces a unique challenge to the field of VLMs by focusing on optical illusions. This dataset differs from traditional VQA datasets by specifically testing the models' comprehension and spatial reasoning abilities through its two main tasks: comprehension and soft localization. These tasks assess VLMs' capability to interpret and localize elements within images that are inherently deceptive, extending beyond basic object recognition to probe deeper into advanced visual processing.
Initial testing on the IllusionVQA dataset highlights significant shortcomings of contemporary VLMs, including GPT-4V. The underperformance compared to near-perfect human results underscores the current gap in VLMs' ability to handle complex visual inputs.These findings point to a pressing need for advancements in VLM architectures and training methods.
Reasons to reject
This paper lacks clear descriptions of the data sources for its optical illusion images, raising concerns about the diversity and representativeness of the dataset. There is no detailed methodology on ensuring data comprehensiveness, which is crucial for a varied visual category like optical illusions. Furthermore, the study does not discuss whether the illusions are representative of real-world scenarios, which would be vital for practical applications of VLMs. Addressing these issues could enhance the dataset’s utility and relevance, providing a stronger benchmark for VLMs in practical environments.