Back

First-24-hour machine learning for 30-day mortality prediction in ICU trauma patients: development in MIMIC-III and cross-database evaluation in MIMIC-IV

Kudrot, N.; Si, Y.; Sanjaya, J.; Pathak, S.; Haghi, M.; Alaei, K.; Placencia, G.; Pishgar, M.

2026-07-30 intensive care and critical care medicine
10.64898/2026.07.28.26359155 medRxiv
Show abstract

ICU trauma patients are clinically heterogeneous, and early mortality risk stratification may support monitoring and resource allocation. We developed machine learning models for 30-day mortality prediction using information recorded during the first 24 hours after ICU admission. In MIMIC-III, six feature configurations were trained using 3,411 patients and compared in a patient-level configuration-selection hold-out subset of 853 patients. The selected XGBoost configuration yielded an area under the precision-recall curve (AUPRC) of 0.556 and an area under the receiver operating characteristic curve (AUROC) of 0.863. For cross-database evaluation, a 228-predictor harmonized XGBoost model was refitted on the complete MIMIC-III cohort and evaluated in 13,747 MIMIC-IV ICU stays without using MIMIC-IV outcomes for model development or recalibration. It achieved an AUPRC of 0.495, an AUROC of 0.825, and a Brier score of 0.109. Calibration was monotonic but showed increasing overprediction at higher predicted risks. First-24-hour clinical information retained predictive value across MIMIC database versions, although internal configuration selection, model differences, same-center provenance, and incomplete feature-mapping documentation limit generalizability and deployment readiness.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.