Improving Clinical Applicability of Heart Failure Readmission Prediction via Automated Feature Engineering
Oloko-Oba, M. O.; Aslam, A.; Echols, M.; Onwuanyi, A.; Idris, M. Y.
Show abstract
Heart failure (HF) readmission prediction models often rely on manually curated, cross-sectional features and show limited discrimination and calibration. We evaluated whether automated feature engineering via Deep Feature Synthesis (DFS) improves the clinical applicability of HF readmission prediction from lon-gitudinal electronic health record data. Using 355,217 HF hospitalizations from a large U.S. safety-net health system (2010-2025), we compared a clinician-curated baseline feature set to DFS-enhanced features and trained identical models for 30-, 60-, and 90-day read-mission. DFS consistently improved gradient-boosted tree performance, increasing AUROC and AUPRC across all horizons, while logistic regression performance declined. At sensitivity-targeted operating points (80%), DFS improved specificity and positive predictive value for boosted trees, reducing false-positive workload. Calibration also improved for boosted trees at all horizons but not for linear models. These results show that automated feature engineering yields deployment-relevant gains that are strongly model-class dependent. Data and Code AvailabilityThis study uses retrospective electronic health record data from a large urban safety-net healthcare system in the United States. Due to patient privacy, institutional restrictions, and data use agreements, the data are not publicly available. An anonymized version of the code used for data processing, feature engineering, model training, and evaluation will be made available upon acceptance of the paper. Institutional Review Board (IRB)This retrospective study was reviewed and approved by an institutional review board. Full IRB details will be provided in the camera-ready version of the paper if accepted.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- High-throughput Phenotyping with Temporal Sequences 93%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 96%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
Similar papers in this journal
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.