Machine learning for elective caesarean section in Bangladesh: validation design, not model choice, determines the performance a deployed model would have
Rony, A. R.; Nahin, K. S. A.; Islam, T.; Asha, A. S.; Hossen, A.
Show abstract
Caesarean section in Bangladesh reached 51.8% of deliveries in 2025, and elective caesarean, meaning caesarean before labour began, reached 31.6%. Risk models built on national household surveys are increasingly proposed for pointing audit toward places where scheduled surgery is outrunning clinical need, but they are usually validated in ways that flatter them. Using the 2025 Bangladesh Multiple Indicator Cluster Survey, we developed four models on 9,538 women (logistic regression, elastic net, random forest, gradient boosting) and ran the same procedure under three validation designs: random five-fold cross-validation; five-fold cross-validation grouped by sampling cluster; and leave-one-division-out cross-validation. We also tested transfer between the 2019 and 2025 rounds and audited subgroup calibration. No model improved on logistic regression by a margin worth acting on: the area under the receiver operating characteristic curve ranged from 0.724 to 0.736 under cluster-grouped validation, a spread of 0.012. Validation design mattered far more than the algorithm. Grouping folds by sampling cluster changed discrimination by at most 0.0004, this survey contributing a median of 3 eligible women per enumeration area. Withholding a whole division cost 0.044 to 0.060, more than 100 times as much, and still cost 0.033 to 0.056 after the strongest predictor, an outcome-derived district rate, was removed from every model. A model fitted to 2019 data lost 0.083 when applied to 2025, and the two rounds agreed only moderately on which predictors mattered (Spearman rank correlation 0.61). Calibration held in every wealth quintile, both residence categories and seven of eight divisions; Sylhet was the exception. Elective caesarean is predictable from routine survey items, but that predictability is local. Cross-validation, including cluster-aware cross-validation, does not measure what a model would do in a district it has never seen; a geographic holdout is the cheapest design that does.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Cohort profile of the ICMR-Stillbirth Pooled India Cohort (ICMR-SPIC): Estimating Prevalence, Analyzing Risk Factors, and Developing Prediction Models for Stillbirths in India 92%
- Birhan Maternal and Child Health cohort: a study protocol 90%
- The effect of prenatal multiple micronutrient supplementation on birth weight in Ethiopia: protocol for a pragmatic cluster-randomised trial 90%
Similar papers in this journal
- Trends and determinants of prelacteal feeding practice in rural Bangladesh from 2004 to 2019: A multivariate decomposition analysis. 90%
- Trends, wealth inequalities and the role of the private sector in caesarean section in the Middle East and North Africa: a repeat cross-sectional analysis of population-based surveys 90%
- Population birth outcomes in 2020 and experiences of expectant mothers during the COVID-19 pandemic: a ‘Born in Wales’ mixed methods study using routine data 90%
Similar papers in this journal
- Better individual-level risk models can improve the targeting and life-saving potential of early-mortality interventions 93%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 91%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 89%
Similar papers in this journal
- Bayesian machine learning enables discovery of risk factors for hepatosplenic multimorbidity related to schistosomiasis 88%
- Evaluating the burden of COVID-19 on hospital resources in Bahia, Brazil: A modelling-based analysis of 14.8 million individuals 88%
- WASH interventions and child diarrhea at the interface of climate and socioeconomic position in Bangladesh 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.