Back

Development of a Machine Learning Model for Predicting In-Hospital Mortality and Analyzing Associated Risk Factors

Liu, J.; He, H.; Wang, Y.; Du, J.; Liang, Y.; Xue, J.; Liang, Y.; Chen, P.; Yang, Q.; Yin, Y.; Wang, G.; Jiang, X.; Deng, Y.

2025-03-01 neurology
10.1101/2025.02.27.25323062 medRxiv
Show abstract

ObjectiveThis study endeavors to construct a machine learning model to forecast in-hospital mortality and dissect associated risk factors, utilizing a vast dataset from multiple hospitals in Chongqing. MethodsWe amassed detailed baseline data encompassing demographics, medical histories, laboratory tests, and imaging indicators from 23,307 ischemic stroke patients. The NIHSS score was derived from admission records, and both in-hospital survival status and causes of death were meticulously documented. Employing the missForest method, we imputed missing values, addressing data imbalance through random oversampling, validated via five-fold cross-validation. The SHAPRFECV technique was instrumental in identifying the most impactful features, steering clear of multicollinearity. A suite of machine learning models, including LR, RF, and KNN, were meticulously tuned using three-fold cross-validation and grid search to optimize hyperparameters. ResultsOur cohort had an average age of 67.347 {+/-} 12.822 years, a baseline NIHSS score of 8.430 {+/-} 3.162, and a 51.186% male predominance, with an in-hospital mortality rate of 6.183%. The Random Forest model excelled with an AUC of 0.940 in the test set, trailed closely by CatBoost at 0.937, LightGBM at 0.930, and XGBoost at 0.929. Notably, CatBoost boasted the highest F1 score of 0.595420 on the test set, with no significant predictive performance disparity between it and the Random Forest model (p = 0.500). ConclusionGrounded in data from four hospitals in Chongqing, our machine learning model, predicated on baseline features, not only streamlines clinical application but also ensures robust predictive efficacy. It provides an in-depth analysis of mortality risk factors, serving as a pivotal reference for clinical decision-making. Future endeavors will concentrate on validating the model within larger-scale, geographically diverse samples, thereby amplifying its applicability and value in clinical practice.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.