Back

Comparative Performance of Machine Learning Models in Predicting Fertility Based on Insights from BDHS Data

Ador, M. Y. H.; Dev Sharma, S. K.; Rukonuzzaman, M.; Chakma, F.; Kamruzzaman, M.

2025-12-17 public and global health
10.64898/2025.12.16.25342342 medRxiv
Show abstract

BackgroundFertility is a social indicator that represents the countrys growth and economic sustainability. The fertility rate in a country signifies the average number of kids that a woman gives birth to throughout her lifetime. The current research is going to use several machine learning models in such a way that they would be capable of detecting the factors that are driving and are responsible for the fertility rate in Bangladesh. MethodsThe data used for this study was obtained from the Bangladesh Demographic Health Survey (BDHS), which was conducted in 2021-22. A variety of machine learning (ML) models and techniques were put into practice including Random Forests (RF), Decision Trees (DT), K-Nearest Neighbors (KNN), Logistic Regression (LR), Support Vector Machines (SVM), XGBoost, LightGBM, Neural Networks (NN), Stacking, and Voting. Along with K-fold cross-validation, Metrics of Accuracy, F1-Score (weighted), Precision (weighted), Recall (weighted), Area under the Receiver Operating Characteristics Curve (AUROC) (weighted), and the weighted average of the Confusion Matrix were applied to the assessment and comparison of the performance of the predictive models. ResultsThis research unveils the discussion on traditional methods and Machine Learning methods, and we found that division, place of residence, religion, and wealth index), mothers education fathers education fathers occupation), mothers occupation, type of toilet facilities), and source of drinking water, contraception use were strongly associated with fertility. With the help of our best identified model Stacking, Voting, and Logistic Regression showed the best results with the highest accuracy (81%), F1-score ([~]78%), and AUC ROC (81%), indicating strong and balanced predictive performance for predicting the determinants influencing the fertility in Bangladesh. Conclusion & RecommendationsAccording to our study Stacking, Voting and Logistics Regression showed better prediction for predicting fertility in Bangladesh. Comparative with analysis with advanced techniques can be done in future work. Moreover, our policy makers and government can take necessary steps by focusing on key determinants influencing fertility.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.