AI-driven fabrication of healthcare survey data: methods, motivations, and ethical implications

ABSTRACT Large language models can now generate synthetic survey data that meet standard reliability and validity thresholds without involving real participants. This study demonstrates how generative AI can replicate the structure of a published healthcare survey to produce highly realistic data that pass conventional psychometric checks, including correlations, loadings, and Cronbach’s alpha. Using Partial Least Squares Structural Equation Modeling (PLS-SEM) via SmartPLS, we show that minimal effort is needed to fabricate plausible datasets. This raises serious ethical concerns, as academic pressure may drive misuse by researchers or students. Standard validation metrics fail to detect such AI-generated responses, creating risks for healthcare research by potentially distorting clinical guidelines and undermining public trust in evidence-based practice. We propose safeguards including AI-generated data audits, dynamic authenticity checks, and ethics training to address this threat. Our findings highlight the urgent need for multi-layered protections to uphold research integrity in an era where artificial intelligence can closely mimic real human data.

Paper

Full text

PDF

AI-driven fabrication of healthcare survey data: methods, motivations, and ethical implications

Semantic Scholar · 2025

Abstract

ABSTRACT Large language models can now generate synthetic survey data that meet standard reliability and validity thresholds without involving real participants. This study demonstrates how generative AI can replicate the structure of a published healthcare survey to produce highly realistic data that pass conventional psychometric checks, including correlations, loadings, and Cronbach’s alpha. Using Partial Least Squares Structural Equation Modeling (PLS-SEM) via SmartPLS, we show that minimal effort is needed to fabricate plausible datasets. This raises serious ethical concerns, as academic pressure may drive misuse by researchers or students. Standard validation metrics fail to detect such AI-generated responses, creating risks for healthcare research by potentially distorting clinical guidelines and undermining public trust in evidence-based practice. We propose safeguards including AI-generated data audits, dynamic authenticity checks, and ethics training to address this threat. Our findings highlight the urgent need for multi-layered protections to uphold research integrity in an era where artificial intelligence can closely mimic real human data.

Similar papers

© 2026 NYSGPT2525 LLC