Back

Machine Learning to Predict Neonatal Mortality Using Public Health Data from Sao Paulo - Brazil

Beluzo, C. E.; Alves, L. C.; Silva, E.; Bresan, R. C.; Arruda, N. M.; Carvalho, T. J. d.

2020-06-22 health informatics
10.1101/2020.06.19.20112953 medRxiv
Show abstract

Infant mortality is one of the most important socioeconomic and health quality indicators in the world. In Brazil, neonatal mortality accounts to 70% of the infant mortality. Despite its importance, neonatal mortality shows increasing signals, which causes concerns about the necessity of efficient and effective methods able to help reducing it. In this paper a new approach is proposed to classify newborns that may be susceptible to neonatal mortality by applying supervised machine learning methods on public health features. The approach is evaluated in a sample of 15,858 records extracted from SPNeoDeath dataset, which were created on this paper, from SINASC and SIM databases from Sao Paulo city (Brazil) for this paper intent. As a results an average AUC of 0.96 was achieved in classifying samples as susceptible to death or not with SVM, XGBoost, Logistic Regression and Random Forests machine learning algorithms. Furthermore the SHAP method was used to understand the features that mostly influenced the algorithms output.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.