Synthetic data have been proposed to facilitate data sharing in privacy-sensitive contexts, including clinical trials. It remains unclear, however, how well original treatment effect estimates can be replicated in synthetic data analyses. Therefore, we synthesized and reanalyzed 128 treatment comparisons from 115 phase 3 randomized oncology trials using sixteen different generative models. Our findings demonstrate that careful methodological choices are essential for drawing valid statistical conclusions from synthetic data analyses. Naive analyses frequently yield falsely significant treatment effects, occurring in up to half of the trials created by deep generative models. Correcting standard errors to reflect the uncertainty inherent in synthetic data generation reduces these false positives, but primarily suffices for trials generated by parametric models. Although this correction entails some power loss, it can be mitigated by increasing the synthetic sample size. Thus, at present, large synthetic trials generated by parametric models and analyzed with corrected standard errors are more likely to preserve inferential utility. Advancing valid statistical inference from synthetic data created by deep generative models remains an important direction for future research.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex