Whom Does Algorithmic Risk Stratification Miss? A Fairness Audit of Machine Learning Targeting for Concurrent Maternal-Child Double Burden of Malnutrition Across 30 Low- and Middle-Income Countries
WU, X.; Zheng, B.
Show abstract
BackgroundConcurrent maternal-child double burden of malnutrition (DBM) affects a growing share of mother-child dyads in low- and middle-income countries (LMICs). Nutrition programmes often use maternal education as an eligibility proxy, but whether algorithmic alternatives would do better--and at what equity cost--has not been directly tested. We evaluated whether machine learning (ML)-based targeting for two concurrent DBM subtypes--overweight mother with stunted or wasted child (Subtype A) and underweight mother with stunted or wasted child (Subtype B)--improves recall over a proxy-based rule while preserving fairness across social strata. MethodsWe pooled Phase 7-8 Demographic and Health Surveys from 30 LMICs (181,636 mother-child dyads). We first estimated subtype-specific social gradients with multilevel logistic regression. We then trained xgboost prediction models with strict label-leakage safeguards and leave-one-country-out cross-validation, and compared ML-based targeting against random and education-based rules at 10%, 20%, and 30% budget constraints. Fairness was audited along six social strata using equalized-odds, demographic-parity, calibration, and predictive-value gaps. A full-India sensitivity analysis (354,691 dyads) assessed robustness to down-sampling. FindingsOverall weighted any-DBM prevalence was 12.52% (Subtype A: 8.21%; Subtype B: 4.31%). Subtype A showed an inverted-U gradient on wealth (adjusted odds ratio peak 1.22 at Richer versus Poorest) and maternal education (peak 1.23 at Primary versus None); Subtype B declined monotonically (wealth: 0.25 at Richest; higher education: 0.49). Mean leave-one-country-out area under the curve was 0.615 for Subtype A and 0.652 for Subtype B. At a 20% budget, ML captured 35.3% of Subtype A cases versus 18.4% for education-based targeting (+92%); for Subtype B the corresponding values were 37.5% and 32.1% (+17%). Equalized-odds gaps reached 0.57 on country income and 0.59 on maternal education; true-positive rates were lowest in the highest-wealth and highest-education strata. Results were stable under the full-India sensitivity analysis. ConclusionsML is useful principally for Subtype A, where the education proxy is no better than random. For Subtype B it mostly changes who gets reached rather than how many, which is a policy choice rather than an accuracy upgrade. The households the algorithm most often misses are not the poor but the rare positives in high-resource strata, which is what a fixed-budget rule ranking on heterogeneous base rates will do. Programmes should decide whether their priority is total capture or the distribution of capture before adopting such a rule. Author SummaryIn many low- and middle-income countries, mothers who are overweight often live in the same household as children who are too short or too thin for their age. Nutrition programmes that try to reach such families have limited resources, so they must choose which households to prioritise. Most programmes use maternal education level as a rough filter, but whether this is actually a good way to find affected families has rarely been tested. We used surveys of 181,636 mother-child pairs from 30 low- and middle-income countries to compare three ways of identifying at-risk households: random selection, selection by low maternal education, and selection by a machine-learning model. Machine learning was much better at finding families where an overweight mother lives with an undernourished child--nearly doubling the capture rate compared with the education rule. For a different combination (underweight mother with an undernourished child), machine learning did not clearly outperform education on total recall; instead it reached different households, mostly shifting attention toward the rural poor. An unexpected finding was that the households the algorithm was most likely to miss were not the poor ones, but the wealthier and better-educated ones, where this type of malnutrition is rarer. This is not bias against the poor--it is what happens when any ranking rule operates under a fixed budget. Programmes that want to reach everyone at risk, regardless of how rare risk is in a given group, may need more than one rule.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Association of Trends in Child Undernutrition and Implementation of the National Rural Health Mission in India: A Nationally Representative Serial Cross-Sectional Study on Data from 1992 – 2015 92%
- The impact of geographic targeting of oral cholera vaccination in sub-Saharan Africa: a modeling study 91%
- The Association Between Cognitive Ability and Body Mass Index: A Sibling-Comparison Analysis in Four Longitudinal Studies 90%
Similar papers in this journal
- Drought, armed conflict and population mortality in Somalia, 2014-2018: a statistical analysis 93%
- Collecting mortality data via mobile phone surveys: a non-inferiority randomized trial in Malawi 93%
- Economic evaluation of participatory women’s groups scaled up by the public health system to improve birth outcomes in Jharkhand, eastern India 93%
Similar papers in this journal
- Effective coverage for maternal health: operationalizing effective coverage cascades for antenatal care and nutrition interventions for pregnant women in seven low- and middle-income countries 95%
- Delays in accessing high-quality care for newborns in East Africa: An analysis of survey data in Malawi, Mozambique, and Tanzania 93%
- Analyzing concordance between MUAC, MUACZ, and WHZ in diagnosing acute malnutrition among children under 5 in Somalia 93%
Similar papers in this journal
- A causal inference approach for estimating effects of non-pharmaceutical interventions during Covid-19 pandemic 93%
- Using Google Health Trends to investigate COVID-19 incidence in Africa 92%
- Effects of trust, risk perception, and health behavior on COVID-19 disease burden: Evidence from a multi-state US survey 91%
Similar papers in this journal
- Changes in stillbirths and child and youth mortality in 2020 and 2021 during the Covid-19 pandemic 92%
- Lifetime risk of maternal near miss morbidity: A novel indicator of maternal health 91%
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.