An Explainable AI Pipeline for Lung CT Analysis: Segmentation, Uncertainty Quantification, and Automated Vision–Language Reporting

Accurate pulmonary infection detection in chest CT is essential to enable timely diagnosis and clinical decisions, particularly amidst large-scale respiratory outbreaks like COVID-19. Traditional interpretation of CT slices by human radiologists is a time-consuming, subjective process, and it is challenging to maintain consistency among observations by different radiologists. Deep learning segmentation models have shown excellent automatic segmentations, yet challenges persist regarding model reliability and uncertainty estimation, while clinical interpretability also remains an open challenge. The proposed framework presents a hybrid architecture that couples a U-Net backbone with Monte Carlo Dropout for pixel-wise uncertainty quantification and Grad-CAM for visual explainability. Furthermore, a BLIP vision–language module is integrated to yield concise clinically relevant descriptions in caption style from the predicted masks and attention maps. Operating at 256 times 256 axial CT slices using T=30 stochastic passes, the proposed framework achieves a Dice score of 85.2%, an IoU of 74.3%, and pixel accuracy of 96.0%, along with a mean per-pixel uncertainty of 0.048. The entire workflow has been wrapped into an intuitive Streamlit application. These results document that this proposed framework delivers accurate predictions, reliable uncertainty information, and visual–language explanations interpretable by humans, for use as a practical and transparent decision-support tool in radiological image analysis.

Paper

Full text

PDF

An Explainable AI Pipeline for Lung CT Analysis: Segmentation, Uncertainty Quantification, and Automated Vision–Language Reporting

Semantic Scholar · 2026

Abstract

Accurate pulmonary infection detection in chest CT is essential to enable timely diagnosis and clinical decisions, particularly amidst large-scale respiratory outbreaks like COVID-19. Traditional interpretation of CT slices by human radiologists is a time-consuming, subjective process, and it is challenging to maintain consistency among observations by different radiologists. Deep learning segmentation models have shown excellent automatic segmentations, yet challenges persist regarding model reliability and uncertainty estimation, while clinical interpretability also remains an open challenge. The proposed framework presents a hybrid architecture that couples a U-Net backbone with Monte Carlo Dropout for pixel-wise uncertainty quantification and Grad-CAM for visual explainability. Furthermore, a BLIP vision–language module is integrated to yield concise clinically relevant descriptions in caption style from the predicted masks and attention maps. Operating at 256 times 256 axial CT slices using T=30 stochastic passes, the proposed framework achieves a Dice score of 85.2%, an IoU of 74.3%, and pixel accuracy of 96.0%, along with a mean per-pixel uncertainty of 0.048. The entire workflow has been wrapped into an intuitive Streamlit application. These results document that this proposed framework delivers accurate predictions, reliable uncertainty information, and visual–language explanations interpretable by humans, for use as a practical and transparent decision-support tool in radiological image analysis.

Similar papers

© 2026 NYSGPT2525 LLC