Contrastive Learning Improves Critical Event Prediction in COVID-19 Patients

Machine Learning (ML) models typically require large-scale, balanced training\ndata to be robust, generalizable, and effective in the context of healthcare.\nThis has been a major issue for developing ML models for the\ncoronavirus-disease 2019 (COVID-19) pandemic where data is highly imbalanced,\nparticularly within electronic health records (EHR) research. Conventional\napproaches in ML use cross-entropy loss (CEL) that often suffers from poor\nmargin classification. For the first time, we show that contrastive loss (CL)\nimproves the performance of CEL especially for imbalanced EHR data and the\nrelated COVID-19 analyses. This study has been approved by the Institutional\nReview Board at the Icahn School of Medicine at Mount Sinai. We use EHR data\nfrom five hospitals within the Mount Sinai Health System (MSHS) to predict\nmortality, intubation, and intensive care unit (ICU) transfer in hospitalized\nCOVID-19 patients over 24 and 48 hour time windows. We train two sequential\narchitectures (RNN and RETAIN) using two loss functions (CEL and CL). Models\nare tested on full sample data set which contain all available data and\nrestricted data set to emulate higher class imbalance.CL models consistently\noutperform CEL models with the restricted data set on these tasks with\ndifferences ranging from 0.04 to 0.15 for AUPRC and 0.05 to 0.1 for AUROC. For\nthe restricted sample, only the CL model maintains proper clustering and is\nable to identify important features, such as pulse oximetry. CL outperforms CEL\nin instances of severe class imbalance, on three EHR outcomes with respect to\nthree performance metrics: predictive power, clustering, and feature\nimportance. We believe that the developed CL framework can be expanded and used\nfor EHR ML work in general.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC