Back

Development and Validation of Machine Learning Models for Predicting Mortality in Hospitalised Systemic Lupus Erythematosus Patients in Dr. Sardjito Hospital, Indonesia Machine Learning Prediction of In-Hospital Mortality in SLE

Paramaiswari, A.; Nugroho, D. B.

2026-05-04 rheumatology
10.64898/2026.05.01.26352268 medRxiv
Show abstract

ObjectivesThis study aimed to develop and validate machine learning models to predict in-hospital mortality among systemic lupus erythematosus (SLE) patients using administrative claims data in a tertiary referral center in Indonesia. MethodsWe conducted a retrospective cohort study of 327 SLE hospital admissions between January 2019 and June 2025. Predictor variables included demographics, hospitalisation characteristics, and the ten most frequent comorbidities. We developed Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost) models. Class imbalance was addressed using the Synthetic Minority Over-sampling Technique. ResultsThe overall in-hospital mortality rate was 7.7%. While models achieved comparable discrimination (Area Under the Curve ~0.71), XGBoost was selected for its superior sensitivity (0.93) compared to Logistic Regression (0.80) and Random Forest (0.97). Feature importance analysis revealed pneumonia as the most significant predictor, followed by acute kidney failure and length of stay. Hypoalbuminemia and hyponatremia were also identified as key prognostic markers. ConclusionsMachine learning models utilising registry-based administrative data effectively stratify mortality risk in hospitalised SLE patients with high sensitivity. The dominance of pneumonia and renal failure as predictors underscores the critical need for aggressive infection control and renal monitoring in this population.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.