Human vs. AI: A Novel Benchmark and a Comparative Study on the Detection of Generated Images and the Impact of Prompts

With the advent of publicly available AI-based text-to-image systems, the\nprocess of creating photorealistic but fully synthetic images has been largely\ndemocratized. This can pose a threat to the public through a simplified spread\nof disinformation. Machine detectors and human media expertise can help to\ndifferentiate between AI-generated (fake) and real images and counteract this\ndanger. Although AI generation models are highly prompt-dependent, the impact\nof the prompt on the fake detection performance has rarely been investigated\nyet. This work therefore examines the influence of the prompt's level of detail\non the detectability of fake images, both with an AI detector and in a user\nstudy. For this purpose, we create a novel dataset, COCOXGEN, which consists of\nreal photos from the COCO dataset as well as images generated with SDXL and\nFooocus using prompts of two standardized lengths. Our user study with 200\nparticipants shows that images generated with longer, more detailed prompts are\ndetected significantly more easily than those generated with short prompts.\nSimilarly, an AI-based detection model achieves better performance on images\ngenerated with longer prompts. However, humans and AI models seem to pay\nattention to different details, as we show in a heat map analysis.\n

Paper

References (24)

Scroll for more · 12 remaining

Similar papers

© 2026 NYSGPT2525 LLC