Reliability and Quality of AI-Generated Information on Newborn Screening Tests: A Comparative Analysis of ChatGPT and Gemini

Background/Objectives: The increasing use of artificial intelligence (AI) chatbots for obtaining health-related information has raised concerns regarding the quality and reliability of the information they provide. This study aimed to compare the quality and reliability of responses generated by ChatGPT free tier (GPT-4o, with GPT-4.1 mini as the fallback model after the usage limit) and Gemini 2.5 Flash (Google, free version) regarding newborn screening tests. Methods: A total of 31 questions were developed based on international and national newborn screening guidelines and were posed to both chatbots. Responses were independently evaluated by two researchers using the DISCERN instrument and the Global Quality Score (GQS), and inter-rater reliability was assessed. Descriptive statistics and non-parametric tests were used to compare chatbot performance. Results: Both evaluators assigned significantly higher DISCERN total scores to Gemini than to ChatGPT free tier. For Evaluator 1, the mean DISCERN scores were 48.8 ± 10.2 for Gemini and 42.4 ± 5.9 for ChatGPT free tier (p < 0.001); for Evaluator 2, the corresponding scores were 53.7 ± 8.9 and 42.4 ± 5.8, respectively (p < 0.001). For the GQS ratings, Evaluator 1 rated Gemini significantly higher than ChatGPT free tier (3.6 ± 0.6 vs. 3.0 ± 0.5, p < 0.001), whereas Evaluator 2 found no statistically significant difference between the two chatbots (3.3 ± 1.3 vs. 3.4 ± 0.8, p = 0.877). Inter-rater reliability for DISCERN scores was excellent for ChatGPT and moderate for Gemini, whereas agreement for GQS ratings was low for both chatbots. Conclusions: Gemini demonstrated consistently higher DISCERN scores than ChatGPT free tier; however, its superiority in overall GQS ratings was not consistently supported across the two evaluators. Neither chatbot consistently achieved the highest levels of information quality. AI chatbots should therefore be considered supplementary sources of health information rather than substitutes for healthcare professionals or official health information resources.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC