Predicting Inpatient Risk of Mortality in Diabetic Patients Using Administrative Data and Machine Learning: An External Validation Study Using SPARCS
Mirza, A. F.; Nwokeji, T. U.
Show abstract
ObjectivesTo evaluate whether machine learning models trained solely on administrative and demographic data can predict inpatient APR Risk of Mortality in diabetic patients. DesignRetrospective cohort study using New York State SPARCS data from 2021 and 2022. SettingNew York Statewide Planning and Research Cooperative System (SPARCS) data from 2021 and 2022. ParticipantsAdult inpatient admissions (age [≥]18) with a diagnosis of diabetes mellitus. Primary outcome measureAPR-DRG Risk of Mortality (ROM), classified as Minor, Moderate, Major, or Extreme. ResultsXGBoost outperformed logistic regression and random forest across all metrics. On the 2022 validation set, XGBoost achieved the highest accuracy (46.5%), macro AUC (0.699), weighted F1-score (0.458), and the lowest Brier score for the Extreme class (0.052). SHAP analysis identified length of stay, age group, and payer type as key predictors. ConclusionsEven without clinical data, administrative features contain non-random signals relevant for mortality risk stratification. These models, especially XGBoost, may help hospitals flag high-risk patients early using routinely available data, aiding triage and planning before labs or vitals are available. Strengths and limitations of this studyO_LIThis study is one of the first to apply machine learning to publicly available SPARCS data to predict APR-DRG Risk of Mortality in diabetic inpatients. C_LIO_LIWe evaluated three models using temporally distinct training and validation cohorts, simulating real-world model deployment across calendar years. C_LIO_LIModel interpretability was addressed using SHAP, providing transparent insights into feature contributions and enabling clinician-facing explanation. C_LIO_LIThe models relied solely on administrative and demographic data, limiting predictive fidelity due to the absence of clinical features such as laboratory values or vital signs. C_LIO_LIRisk of Mortality labels were derived from APR-DRG software and may be influenced by coding practices rather than objective clinical outcomes. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 96%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 95%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 95%
Similar papers in this journal
- Development and validation of a machine learning model for predicting illness trajectory and hospital resource utilization of COVID-19 hospitalized patients - a nationwide study 94%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
Similar papers in this journal
- Predicting mortality in SARS-COV-2 (COVID-19) positive patients in the inpatient setting using a Novel Deep Neural Network 93%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 93%
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 93%
Similar papers in this journal
- An Online Risk Calculator for Rapid Prediction of In-hospital Mortality from COVID-19 Infection 94%
- A Machine Learning-Based Prediction of Hospital Mortality in Mechanically Ventilated ICU Patients 94%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 94%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.