Cross-Continental Transfer of Perioperative Mortality Prediction: Intraoperative Features Generalize Where Preoperative Features Fail
Purkayastha, D. S.
Show abstract
BackgroundMachine learning models for perioperative mortality prediction show strong internal discrimination, yet external validation--particularly across continents--remains rare. Whether intraoperative vital sign features, which improve internal performance by 5-10%, transfer across populations is unknown. Furthermore, aggregate discrimination metrics may overstate clinical utility through Simpsons paradox if models separate risk strata without discriminating within them. MethodsWe conducted a bidirectional cross-continental external validation study using publicly available datasets from Korea (INSPIRE, n=127,413 surgeries, 1,387 deaths) and the United States (MOVER, n=57,545 surgeries, 823 deaths). Eight machine learning models (XGBoost and logistic regression with preoperative-only or preoperative-plus-intraoperative features) were trained on each dataset and validated on the other. We assessed dis-crimination (AUC-ROC), calibration, clinical utility (decision curve analysis), and within-stratum performance to detect Simpsons paradox. ResultsAll models achieved clinically useful external discrimination (AUC >0.70). The best-performing model (XGB-INS-B) achieved external AUC of 0.895, representing a 4.1% improvement over internal performance. Intraoperative models showed higher mean external AUC than preoperative models (0.82 versus 0.79). Models trained on the diverse Korean population showed 2.6-fold less degradation than those trained on the concentrated US high-acuity population (5.4% versus 13.9%). Critically, preoperative models exhibited Simpsons paradox: one model achieved acceptable overall AUC (0.756) while performing at near-random levels within both ASA strata (0.597 and 0.584). Intraoperative models maintained within-stratum discrimination (0.71-0.74). All models required Platt scaling recalibration. ConclusionsIntraoperative vital sign features provide population-independent prognostic information that transfers across continents, while preoperative models achieve apparent discrimination through risk stratum separation. For global deployment of perioperative prediction models, real-time physiological monitoring should be prioritized, and stratified validation is essential to detect Simpsons paradox.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 95%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 93%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 92%
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 89%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 88%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 92%
- White Blood Cell and Platelet Dynamics Define Human Inflammatory Recovery 92%
- Integration of clinical characteristics, lab tests and a deep learning CT scan analysis to predict severity of hospitalized COVID-19 patients 92%
Similar papers in this journal
- Using patient biomarker time series to determine mortality risk in hospitalised COVID-19 patients: a comparative analysis across two New York hospitals 92%
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 92%
- Prognostic pan-cancer and single-cancer models: A large-scale analysis using a real-world clinico-genomic database 90%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 93%
- Personalized survival probabilities for SARS-CoV-2 positive patients by explainable machine learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.