Back

Traditional Machine Learning Outperforms Automated Machine Learning for Postpartum Readmission Prediction: A Comprehensive Performance and Health-Economic Analysis

Crabtree, L.; Wakefield, C.; Gheorghe, C. P.; Frasch, M. G.

2025-12-18 health systems and quality improvement
10.64898/2025.12.16.25342440 medRxiv
Show abstract

ObjectiveAutomated machine learning (AutoML) promises to democratize predictive modeling in healthcare by automating algorithm selection and hyperparameter optimization. Our objective was to compare the performance, clinical utility, and health-economic implications of traditional machine learning versus AutoML frame-works for predicting 14-day postpartum readmission using a large national cohort. MethodsWe analyzed data from 8,774 participants in the nuMoM2b (Nulli-parous Pregnancy Outcomes Study) with complete readmission data. Recursive feature elimination selected 10 sociodemographic predictors from 18 candidates. We applied SMOTE oversampling to the training set for class balancing. Three traditional ML algorithms (logistic regression, random forest, gradient boosting) were compared with FLAML, a leading AutoML framework. Models were evaluated on discrimination (ROC-AUC, PR-AUC), clinical utility (sensitivity/specificity), and computational efficiency using a 70/30 stratified train-test split with bootstrap confidence intervals. ResultsAmong 8,774 participants, 154 (1.8%) experienced 14-day readmission. In a held-out test set (n=2,633; 46 readmissions), logistic regression achieved the high-est ROC-AUC (0.569 [95% CI: 0.483-0.655]) and was the only model with clinically meaningful sensitivity (34.8% [20.5-48.6%]) at the default threshold, correctly identifying 16 of 46 readmissions. FLAML achieved near-chance discrimination (ROC-AUC: 0.500 [0.449-0.552]) with near-zero sensitivity (2.2%). Stacking and calibrated soft-voting ensembles did not improve over logistic regression alone (ROC-AUC: 0.496 and 0.555, respectively). Threshold optimization substantially improved screening performance: lowering the logistic regression threshold to 0.35 increased sensitivity to 82.6% (38/46 readmissions), though at the cost of flagging 76.5% of patients. Health-economic analysis showed that cost-effectiveness requires low-cost interventions (<$59/flagged patient). ConclusionsTraditional logistic regression outperformed AutoML and ensemble methods for postpartum readmission prediction. Threshold optimization, rather than model complexity, provided the largest gains in screening sensitivity. However, all approaches showed modest discrimination using sociodemographic variables alone, establishing a baseline for future models incorporating clinical features. HighlightsO_LITraditional logistic regression outperformed AutoML (FLAML) and ensemble methods (stacking, soft voting), uniquely identifying high-risk patients and demonstrating that more complex algorithms do not guarantee clinical utility for rare outcomes. C_LIO_LIThreshold optimization transformed clinical utility: lowering the logistic regression threshold from 0.50 to 0.35 increased sensitivity from 34.8% to 82.6% (38/46 readmissions captured), enabling use as a screening tool despite modest overall discrimination (ROC-AUC: 0.569). C_LIO_LIHealth-economic value depends critically on intervention cost: conventional enhanced discharge programs are not cost-effective at the models PPV (2.1%), but low-cost triage strategies (<$59/flagged patient) achieve positive ROI. C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.