Observation-process features are associated with larger domain shift in sepsis mortality prediction: a cross-database evaluation using MIMIC-IV and eICU-CRD
Yamamoto, R.; Wu, F.; Sprehe, L. K.; Abeer, A.; Celi, L. A.; Tohyama, T.
Show abstract
Clinical prediction models for sepsis frequently degrade when applied outside the development setting. Electronic health record data encode not only patient physiology but also observation processes such as measurement timing and frequency, which may be predictive within a site but unstable across sites. The contribution of these observation-process features to cross-site performance degradation has not been quantified. In this retrospective cohort study, we developed models for in-hospital mortality in adult intensive care unit (ICU) patients meeting Sepsis-3 criteria using Medical Information Mart for Intensive Care IV (MIMIC-IV) (n = 30,218; 16.3% mortality) and externally validated them in eICU Collaborative Research Database (eICU-CRD) (n = 31,403; 13.9% mortality). We compared seven prespecified model specifications representing physiologic summary strategies (a single aggregate severity score, most recent values, extreme values, and within-window variability), each evaluated with and without measurement counts as observation-process features. Models were fit using logistic regression and gradient-boosted trees. Internally, discrimination improved with more detailed physiologic summaries and measurement counts (logistic regression area under the receiver operating characteristic curve [AUROC] from 0.819 to 0.834). In external validation, performance drops were larger for specifications using more complex physiologic representations. Adding measurement counts was associated with larger domain shift (AUROC change, -0.047 versus -0.082 with counts in logistic regression). External calibration deteriorated progressively, with calibration slopes decreasing from 1.007 for the simplest model to 0.417 for the most complex specification in logistic regression. Gradient-boosted trees showed smaller incremental degradation from measurement counts but still exhibited domain shift in complex specifications. Inclusion of observation-process features in sepsis mortality prediction models was associated with improved internal discrimination but worse external calibration and transportability. These findings highlight that feature engineering decisions involve a tradeoff between internal performance and external generalizability, and that calibration assessment provides the most sensitive indicator of reduced transportability.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A comprehensive ML-based Respiratory Monitoring System for Physiological Monitoring & Resource Planning in the ICU 94%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 94%
Similar papers in this journal
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 96%
- Using patient biomarker time series to determine mortality risk in hospitalised COVID-19 patients: a comparative analysis across two New York hospitals 93%
- SOFA score performs worse than age for predicting mortality in patients with COVID-19 93%
Similar papers in this journal
- Predicting bloodstream infection outcome using machine learning 96%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 96%
- Developing And Validating COVID-19 Adverse Outcome Risk Prediction Models From A Bi-National European Cohort Of 5594 Patients 94%
Similar papers in this journal
- SWIFT: A Deep Learning Approach to Prediction of Hypoxemic Events in Critically-Ill Patients Using SpO 2 Waveform Prediction 93%
- Contrasting factors associated with COVID-19-related ICU admission and death outcomes in hospitalised patients by means of Shapley values 93%
- Predicting the causative pathogen among children with pneumonia using a causal Bayesian network 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.