Reducing Inequalities Using an Unbiased Machine Learning Approach to Identify Births with the Highest Risk of Preventable Neonatal Deaths.
Ramos, A. P.; Caldieraro, F.; Nascimento, M. L.; Saldanha, R.
Show abstract
BackgroundDespite contemporaneous declines in neonatal mortality, recent studies show the existence of left-behind populations that continue to have higher mortality rates than the national averages. Additionally, many of these deaths are from preventable causes. This reality creates the need for more precise methods to identify high-risk births, allowing policymakers to target them more effectively. This study fills this gap by developing unbiased machine-learning approaches to more accurately identify births with a high risk of neonatal deaths from preventable causes. MethodsWe link administrative databases from the Brazilian health ministry to obtain birth and death records in the country from 2015 to 2017. The final dataset comprises 8,797,968 births, of which 59,615 newborns died before reaching 28 days alive (neonatal deaths). These neonatal deaths are categorized into preventable deaths (42,290) and non-preventable deaths (17,325). Our analysis identifies the death risk of the former group, as they are amenable to policy interventions. We train six machine-learning algorithms, test their performance on unseen data, and evaluate them using a new policy-oriented metric. To avoid biased policy recommendations, we also investigate how our approach impacts disadvantaged populations. ResultsXGBoost was the best-performing algorithm for our task, with the 5% of births identified as highest risk by the model accounting for over 85% of the observed deaths. Furthermore, the risk predictions exhibit no statistical differences in the proportion of actual preventable deaths from disadvantaged populations, defined by race, education, marital status, and maternal age. These results are similar for other threshold levels. ConclusionsWe show that, by using publicly available administrative data sets and ML methods, it is possible to identify the births with the highest risk of preventable deaths with a high degree of accuracy. This is useful for policymakers as they can target health interventions to those who need them the most and where they can be effective without producing bias against disadvantaged populations. Overall, our approach can guide policymakers in reducing neonatal mortality rates and their health inequalities. Finally, it can be adapted for use in other developing countries.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Better individual-level risk models can improve the targeting and life-saving potential of early-mortality interventions 96%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 94%
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 94%
Similar papers in this journal
- Identification of high-risk COVID-19 patients using machine learning 93%
- A Bayesian Susceptible-Infectious-Hospitalized-Ventilated-Recovered Model to Predict Demand for COVID-19 Inpatient Care in a Large Healthcare System 92%
- Public participation in crisis policymaking. How 30,000 Dutch citizens advised their government on relaxing COVID-19 lockdown measures 92%
Similar papers in this journal
- Comparing methods to predict baseline mortality for excess mortality calculations 93%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 93%
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 92%
Similar papers in this journal
- Ensemble Machine Learning Modeling for the Prediction of Artemisinin Resistance in Malaria 90%
- Meta-Signer: Metagenomic Signature Identifier based on Rank Aggregation of Features 89%
- Child health, nutrition and gut microbiota development during the first two years of life; study protocol of a prospective cohort study from the Khyber Pakhtunkhwa, Pakistan 89%
Similar papers in this journal
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 93%
- Network meta-analysis and random walks 92%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.