Traditional Machine Learning Outperforms Automated Machine Learning for Postpartum Readmission Prediction: A Comprehensive Performance and Health-Economic Analysis
Crabtree, L.; Wakefield, C.; Gheorghe, C. P.; Frasch, M. G.
Show abstract
ObjectiveAutomated machine learning (AutoML) promises to democratize predictive modeling in healthcare by automating algorithm selection and hyperparameter optimization. Our objective was to compare the performance, clinical utility, and health-economic implications of traditional machine learning versus AutoML frame-works for predicting 14-day postpartum readmission using a large national cohort. MethodsWe analyzed data from 8,774 participants in the nuMoM2b (Nulli-parous Pregnancy Outcomes Study) with complete readmission data. Recursive feature elimination selected 10 sociodemographic predictors from 18 candidates. We applied SMOTE oversampling to the training set for class balancing. Three traditional ML algorithms (logistic regression, random forest, gradient boosting) were compared with FLAML, a leading AutoML framework. Models were evaluated on discrimination (ROC-AUC, PR-AUC), clinical utility (sensitivity/specificity), and computational efficiency using a 70/30 stratified train-test split with bootstrap confidence intervals. ResultsAmong 8,774 participants, 154 (1.8%) experienced 14-day readmission. In a held-out test set (n=2,633; 46 readmissions), logistic regression achieved the high-est ROC-AUC (0.569 [95% CI: 0.483-0.655]) and was the only model with clinically meaningful sensitivity (34.8% [20.5-48.6%]) at the default threshold, correctly identifying 16 of 46 readmissions. FLAML achieved near-chance discrimination (ROC-AUC: 0.500 [0.449-0.552]) with near-zero sensitivity (2.2%). Stacking and calibrated soft-voting ensembles did not improve over logistic regression alone (ROC-AUC: 0.496 and 0.555, respectively). Threshold optimization substantially improved screening performance: lowering the logistic regression threshold to 0.35 increased sensitivity to 82.6% (38/46 readmissions), though at the cost of flagging 76.5% of patients. Health-economic analysis showed that cost-effectiveness requires low-cost interventions (<$59/flagged patient). ConclusionsTraditional logistic regression outperformed AutoML and ensemble methods for postpartum readmission prediction. Threshold optimization, rather than model complexity, provided the largest gains in screening sensitivity. However, all approaches showed modest discrimination using sociodemographic variables alone, establishing a baseline for future models incorporating clinical features. HighlightsO_LITraditional logistic regression outperformed AutoML (FLAML) and ensemble methods (stacking, soft voting), uniquely identifying high-risk patients and demonstrating that more complex algorithms do not guarantee clinical utility for rare outcomes. C_LIO_LIThreshold optimization transformed clinical utility: lowering the logistic regression threshold from 0.50 to 0.35 increased sensitivity from 34.8% to 82.6% (38/46 readmissions captured), enabling use as a screening tool despite modest overall discrimination (ROC-AUC: 0.569). C_LIO_LIHealth-economic value depends critically on intervention cost: conventional enhanced discharge programs are not cost-effective at the models PPV (2.1%), but low-cost triage strategies (<$59/flagged patient) achieve positive ROI. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 93%
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.