Development of the Short Hospitalization Predictor (SHoP) Machine Learning Model Across Two Hospitals
Leuchter, R. K.; Salari, V.; Gabel, E.
Show abstract
ObjectiveTo develop and evaluate an open-source machine learning (ML) models for predicting hospital short stays (length of stay [LOS] under 48 and 72 hours) exclusively using data available at the time of ED admission, with a novel application of target encoding diagnostic codes. Materials and MethodsWe trained two ML algorithms (Random Forest and XGBoost) on electronic health record (EHR) data from two hospitals to predict hospital short stays. We employed an innovative weighted target encoding method that converted categorical International Classification of Disease (ICD- 10) codes into numeric representations of their probabilistic contribution to LOS. We measured area under the receiver operating characteristic curve (AUC) for correctly predicting LOS under 48 or 72 hours, which we compared to logistic regression. ResultsThe final sample included 8,693 adult patients admitted to an internal medicine service. Random Forest models achieved the highest performance for predicting LOS under 48 hours (AUROC=0.96, 95% CI 0.95-0.97; accuracy=91%) and under 72 hours (AUROC=0.94, 95% CI 0.93-0.95; accuracy=88%). These models outperformed logistic regression using the same features (48-hour AUROC=0.57, 95% CI 0.54-0.59 and accuracy=70%; 72-hour AUROC=0.59, 95% CI 0.57-0.61 and accuracy=56%). DiscussionLeveraging an innovative target encoding method, the Short Hospitalization Prediction (SHoP) model substantially outperforms previous ML approaches in accurately predicting LOS under both 48 and 72 hours using only ED pre-admission data (AUC 0.94-0.96). ConclusionThe technical innovation and predictive capability of the SHoP model enables powerful, real-time applications for optimizing patient flow and hospital resource utilization by identifying potentially divertible admissions while patients are still in the ED.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 93%
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 93%
- Modeling physician variability to prioritize relevant medical record information 93%
Similar papers in this journal
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 96%
- Consistency of performance of adverse outcome prediction models for hospitalized COVID-19 patients 95%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
Similar papers in this journal
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 96%
- Predicting patients with false negative SARS-CoV-2 testing at hospital admission: A retrospective multi-center study 95%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 94%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 97%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 95%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 95%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.