Comparison of machine-learning and logistic regression models to predict 30-day unplanned readmission: a development and validation study
Iwagami, M.; Inokuchi, R.; Kawakami, E.; Yamada, T.; Goto, A.; Kuno, T.; Hashimoto, Y.; Michihata, N.; Goto, T.; Shinozaki, T.; Sun, Y.; Taniguchi, Y.; Komiyama, J.; Uda, K.; Abe, T.; Tamiya, N.
Show abstract
We compared the predictive performance of gradient-boosted decision tree (GBDT), random forest (RF), deep neural network (DNN), and logistic regression (LR) with the least absolute shrinkage and selection operator (LASSO) for 30-day unplanned readmission, according to the number of predictor variables and presence/absence of blood-test results. We used electronic health records of patients discharged alive from 38 hospitals in 2015-2017 for derivation (n=339,513) and in 2018 for validation (n=118,074), including basic characteristics (age, sex, admission diagnosis category, number of hospitalizations in the past year, discharge location), diagnosis, surgery, procedure, and drug codes, and blood-test results. We created six patterns of datasets having different numbers of binary variables (that [≥]5% or [≥]1% of patients or [≥]10 patients had) with and without blood-test results. For the dataset with the smallest number of variables (102), the c-statistic was highest for GBDT (0.740), followed by RF (0.734), LR-LASSO (0.720), and DNN (0.664). For the dataset with the largest number of variables (1543), the c-statistic was highest for GBDT (0.764), followed by LR-LASSO (0.755), RF (0.751), and DNN (0.720). We found that GBDT generally outperformed LR-LASSO, but the difference became smaller when the number of variables was increased and blood-test results were used.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Medium-term impacts of the waves of the COVID-19 epidemic on treatments for non-COVID-19 patients in intensive care units: a retrospective cohort study in Japan 96%
- Regional performance variation in external validation of four prediction models for severity of COVID-19 at hospital admission: An observational multi-centre cohort study 95%
- Factors Associated With Regional Differences in Healthcare Quality for Patients With Acute Myocardial Infarction in Japan 95%
Similar papers in this journal
Similar papers in this journal
- Nonspecific blood tests as proxies for COVID-19 hospitalization: are there plausible associations after excluding noisy predictors? 94%
- Development and Validation of the Patient History COVID-19 (PH-Covid19) Scoring System: A Multivariable Prediction Model of Death in Mexican Patients with COVID-19 93%
- Excess Mortality in the United States During the First Three Months of the COVID-19 Pandemic 92%
Similar papers in this journal
- Association of age at menarche and menopause, reproductive lifespan, and stroke among Chinese women: Results from a national cohort study 90%
- When will the battle against novel coronavirus end in Wuhan: a SEIR modeling analysis 89%
- Comparison of excess deaths and laboratory-confirmed COVID-19 deaths during a large Omicron epidemic in 2022 in Hong Kong 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.