Evaluating and Improving the Effectiveness of Synthetic Chest X-Rays for Medical Image Analysis
In this work, we explore best-practice approaches for generating synthetic chest X-ray images and augmenting medical imaging datasets to optimize the performance of deep learning models in downstream tasks like classification and segmentation. We utilize a latent diffusion model to condition the generation of synthetic chest X-rays on text prompts and/or segmentation masks. We explore methods such as using a proxy model and incorporating radiologist feedback to improve the quality of synthetic data. These synthetic images are generated from relevant disease information or geometrically-transformed segmentation masks and added to ground truth training set images from the CheXpert [9], CANDID-PTX[4], SIIM [19], and RSNA Pneumonia [17] in order to measure improvements in classification and segmentation model performance on the test sets. F1 and Dice scores are used to evaluate classification and segmentation, respectively. Across all experiments, the synthetic data we generate results in a maximum mean classification F1 score improvement of 0.15 (CI: 0.10, 0.20; P=0.0031) compared to using only real data. For segmentation, the maximum Dice score improvement is 0.14 (CI: 0.11, 0.18; P=0.0064). We find that best practices for generating synthetic chest Xray images for downstream tasks include conditioning on single-disease labels or geometrically-transformed segmentation masks, as well as potentially using proxy modeling to fine-tune such generations. “Based on these chest X-ray reports, please write a five-word caption with the main finding. Don‘t make comparisons with previous studies, so do not use words such as ‘unchanged’, ‘improved’, ‘worsened’, ‘no change’, ‘increased’, ‘decreased’, etc. in the caption. Don‘t use commas or quotation marks in the caption. If it is normal or no problems are detected, just return ‘Normal’ as the caption.”
Paper
References (24)
Scroll for more · 12 remaining