Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study

Machine learning models for clinical prediction tasks, such as in-hospital mortality and sepsis onset, routinely achieve high AUROC scores. However, AUROC measures ranking performance rather than clinical sensibility. A model may rank patients correctly overall while predicting a lower mortality risk when a patient's SOFA score worsens, contradicting established medical knowledge. This paper pr…

Paper

Similar papers

© 2026 NYSGPT2525 LLC