Class imbalance correction in artificial intelligence models leads to miscalibrated clinical predictions: a real-world evaluation
Roesler, M. W.; Wells, C.; Schamberg, G.; Gao, J.; Harrison, E.; O'Grady, G.; Varghese, C.
Show abstract
BackgroundPredictive models employing machine learning algorithms are increasingly being used in clinical decision making, and improperly calibrated models can result in systematic harm. We sought to investigate the impact of class imbalance correction, a commonly applied preprocessing step in machine learning model development, on calibration and modelled clinical decision making in a large real-world context. MethodsA histogram boosted gradient classifier was trained on a highly imbalanced national dataset of >1.8 million patients undergoing surgery, to predict the risk of 90-day mortality and complications after surgery. Class imbalance correction strategies including random oversampling, synthetic minority oversampling technique, random under-sampling, and cost-sensitive learning were compared to the natural distribution ( natural). Models were tested and compared with classification metrics, calibration plots, decision curve analysis, and simulated clinical impact analysis. ResultsThe natural model demonstrated high performance (AUROC 0.94, 95% CI 0.94-0.95 for mortality; 0.84, 95% CI 0.84-0.85 for complications) and calibration (log loss 0.05, 95% CI 0.04-0.05 for mortality; 0.23, 95% CI 0.23-0.24 for complications). Class imbalance mitigation (CSL, ROS, RUS, and SMOTE) did not improve AUROC or AUPRC but increased recall and F1 scores at the expense of precision and accuracy. However, these methods severely compromised model calibration, leading to significant over-prediction of risks (up to a 62.8 % increase) as further evidenced by increased log loss across all mitigation techniques. Decision curve analysis and clinical scenario testing confirmed that the natural model provided the highest net benefit. ConclusionClass imbalance correction methods result in significant miscalibration, leading to possible harm when used for clinical decision making.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 97%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 96%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 95%
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 95%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 95%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 95%
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 92%
- Descriptive and prognostic value of a computational model of metastasis in high-risk neuroblastoma 91%
Similar papers in this journal
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 95%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 94%
- Predicting the physiological effects of multiple drugs using electronic health record 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.