Assessment of the Performance of ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini 2.0 Flash on Image-Based Ocular Oncology and Pathology Questions: A Cross-Sectional Research
Objective: This study aimed to evaluate the diagnostic performance of 3 flagship models from 3 different companies, Chat Generative Pre-trained Transformer-4 Omni (ChatGPT-4o), Claude 3.5 Sonnet, and Gemini 2.0 Flash, on image-based questions in ocular oncology and pathology to investigate potential differences between these models, and their clinical utility. Material and Methods: Fifty multiple-choice, imagebased questions were randomly selected from 312 questions in the field of ocular oncology and pathology from the OphthoQuestions (www.ophthoquestions.com) database. The answers given to the questions were compared with the answer key and recorded as correct or incorrect. ChatGPT-4o, Claude 3.5 Sonnet and Gemini 2.0 Flash models, which have the ability to process images in large language models (LLMs), were included in the study. Cochran's Q test was applied to compare the performance of the 3 LLMs and McNemar's test was used in pairwise comparisons. Results: There was a statistically significant difference between all 3 LLMs (p=0.001, Cochran's Q test). Claude 3.5 sonnet showed the highest accuracy by correctly identifying 84% of the questions, followed by ChatGPT-4o with 80% and Gemini 2.0 with 62%. In the pairwise comparisons, Claude 3.5 sonnet and ChatGPT-4o were found to be statistically superior to Gemini 2.0 Flash model (p=0.002, p=0.004, respectively). There was no significant difference between Claude 3.5 and ChatGPT-4o (p=0.727, McNemar test). Conclusion: Our results indicate that Claude 3.5 Sonnet and GPT-4o outperform Gemini 2.0 Flash in diagnostic accuracy for ocular oncology and pathology. While LLMs show promise in this field, they require evaluation with larger datasets, and their accuracy must be improved before they can be clinically implemented.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex