Back

Development and External Validation of a High-Precision Model for Predicting ICU Admission from Emergency Department Triage

Nguyen, N. T.; Chu, A. L.; Dash, D.

2025-07-23 health informatics
10.1101/2025.07.22.25332000 medRxiv
Show abstract

ObjectiveTo develop, internally evaluate, and externally validate a machine-learning (ML) model predicting intensive care unit (ICU) admission or death using information available solely at emergency department (ED) triage. Performance was primarily assessed by area under the precision-recall curve (AUPRC) to address severe class imbalance. MethodsWe trained an XGBoost classifier on the Medical Information Mart for Intensive Care IV (MIMIC-IV) dataset. Positive outcomes were ICU admission or death within 6 hours of arrival. Features included vital signs, engineered physiological measures, clinician-assigned acuity, demographics, chief complaint, and home medications. Model performance was internally evaluated through group-stratified five-fold cross-validation and externally validated on the Multimodal Clinical Monitoring in the Emergency Department (MC-MED) dataset. ResultsIn the internal validation (MIMIC-IV, 350,241 visits; 11,745 ICU/death), the model achieved an AUPRC of 0.736 (95% CI: 0.728-0.743), AUROC of 0.966 (95% CI: 0.965-0.968), and accuracy of 0.936 (95% CI: 0.936-0.938). On external validation (MC-MED, 42,624 visits; 1,503 ICU/death), the model retained robust performance with an AUPRC of 0.602 (95% CI: 0.578-0.624, 0.134 decrease), AUROC of 0.949 (95% CI: 0.944-0.955, 0.017 decrease), and accuracy of 0.928 (95% CI: 0.927-0.932, 0.007 decrease), demonstrating promising generalizability despite institutional, temporal, and patient demographic differences. ConclusionsThis study presents one of the first triage ML models externally validated on a distinct ED cohort, achieving a new benchmark for AUPRC in flagging critically ill patients within minutes. Future directions include multi-site training and validation to further enhance real-world generalizability and clinical applicability.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.