An Explainable AI Pipeline for Lung CT Analysis: Segmentation, Uncertainty Quantification, and Automated Vision–Language Reporting
Accurate pulmonary infection detection in chest CT is essential to enable timely diagnosis and clinical decisions, particularly amidst large-scale respiratory outbreaks like COVID-19. Traditional interpretation of CT slices by human radiologists is a time-consuming, subjective process, and it is challenging to maintain consistency among observations by different radiologists. Deep learning segmentation models have shown excellent automatic segmentations, yet challenges persist regarding model reliability and uncertainty estimation, while clinical interpretability also remains an open challenge. The proposed framework presents a hybrid architecture that couples a U-Net backbone with Monte Carlo Dropout for pixel-wise uncertainty quantification and Grad-CAM for visual explainability. Furthermore, a BLIP vision–language module is integrated to yield concise clinically relevant descriptions in caption style from the predicted masks and attention maps. Operating at 256 times 256 axial CT slices using T=30 stochastic passes, the proposed framework achieves a Dice score of 85.2%, an IoU of 74.3%, and pixel accuracy of 96.0%, along with a mean per-pixel uncertainty of 0.048. The entire workflow has been wrapped into an intuitive Streamlit application. These results document that this proposed framework delivers accurate predictions, reliable uncertainty information, and visual–language explanations interpretable by humans, for use as a practical and transparent decision-support tool in radiological image analysis.
Paper
Full text
An Explainable AI Pipeline for Lung CT Analysis: Segmentation, Uncertainty Quantification, and Automated Vision–Language Reporting
Semantic Scholar · 2026
Abstract
Accurate pulmonary infection detection in chest CT is essential to enable timely diagnosis and clinical decisions, particularly amidst large-scale respiratory outbreaks like COVID-19. Traditional interpretation of CT slices by human radiologists is a time-consuming, subjective process, and it is challenging to maintain consistency among observations by different radiologists. Deep learning segmentation models have shown excellent automatic segmentations, yet challenges persist regarding model reliability and uncertainty estimation, while clinical interpretability also remains an open challenge. The proposed framework presents a hybrid architecture that couples a U-Net backbone with Monte Carlo Dropout for pixel-wise uncertainty quantification and Grad-CAM for visual explainability. Furthermore, a BLIP vision–language module is integrated to yield concise clinically relevant descriptions in caption style from the predicted masks and attention maps. Operating at 256 times 256 axial CT slices using T=30 stochastic passes, the proposed framework achieves a Dice score of 85.2%, an IoU of 74.3%, and pixel accuracy of 96.0%, along with a mean per-pixel uncertainty of 0.048. The entire workflow has been wrapped into an intuitive Streamlit application. These results document that this proposed framework delivers accurate predictions, reliable uncertainty information, and visual–language explanations interpretable by humans, for use as a practical and transparent decision-support tool in radiological image analysis.