Back

Predicting county-level diagnosed diabetes prevalence in the United States using explainable gradient boosting and geographic interpretation

Yahaya, Y.; Khan, S.; Rani Saha, P.; Meia, M. A. A.

2026-06-26 endocrinology
10.64898/2026.06.23.26356400 medRxiv
Show abstract

Diagnosed diabetes affects approximately 38.4 million Americans, but its burden is not evenly distributed across U.S. counties. Existing machine-learning studies have mainly focused on individual risk prediction using biometric, clinical, or survey variables. These approaches are less suited to explaining why diagnosed diabetes prevalence differs geographically across counties. We developed an explainable gradient-boosting framework for predicting county-level diagnosed diabetes prevalence across 2,957 U.S. counties using an ecological cross-sectional design. The analysis integrated food-environment, socioeconomic, occupational, demographic, health-behavior, and clinical indicators from five public data sources. Four regression models were compared: Elastic Net, Random Forest, XGBoost, and LightGBM. LightGBM was selected as the primary model based on validation-set RMSE and interpreted using SHAP TreeExplainer. The validation-selected LightGBM model achieved a held-out test RMSE of 0.423 percentage points, R{superscript 2} = 0.964, and MAPE = 2.76%. Although XGBoost achieved a lower test RMSE of 0.399 and R{superscript 2} = 0.968, it was retained as a secondary benchmark because primary-model selection was based only on validation performance. A sensitivity model using only structural and contextual predictors, and excluding CDC PLACES health-behavior and clinical covariates, retained substantial predictive performance (R{superscript 2} = 0.827). Poverty rate was the most frequent dominant positive structural SHAP contributor nationally (n = 772 counties, 26.1%), followed by food insecurity rate (n = 707, 23.9%), Supplemental Nutrition Assistance Program (SNAP) participation rate (n = 316, 10.7%), unemployment rate (n = 224, 7.6%), and median household income (n = 178, 6.0%). Residual Morans I decreased from 0.665 to 0.069 after model fitting. Explainable machine learning using public county-level data can characterize geographic variation in diagnosed diabetes prevalence. County-level SHAP maps may support local hypothesis generation, but should be interpreted as explanations of model predictions rather than causal effects.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 14%
13.1%
2
BMC Medical Research Methodology
47 papers in training set
Top 0.1%
11.0%
3
PLOS Global Public Health
344 papers in training set
Top 2%
8.2%
4
JMIR Public Health and Surveillance
45 papers in training set
Top 0.1%
7.5%
5
Communications Medicine
113 papers in training set
Top 0.8%
3.6%
6
Scientific Reports
3612 papers in training set
Top 31%
3.4%
7
Nature Communications
5641 papers in training set
Top 34%
3.3%
50% of probability mass above
8
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.9%
3.3%
9
Biology Methods and Protocols
61 papers in training set
Top 0.4%
2.5%
10
Diabetes, Obesity and Metabolism
22 papers in training set
Top 0.4%
2.5%
11
eLife
5828 papers in training set
Top 40%
2.5%
12
Current Developments in Nutrition
15 papers in training set
Top 0.3%
2.1%
13
American Journal of Epidemiology
67 papers in training set
Top 0.6%
2.0%
14
The American Journal of Tropical Medicine and Hygiene
68 papers in training set
Top 0.9%
1.8%
15
The American Journal of Clinical Nutrition
19 papers in training set
Top 0.2%
1.8%
16
eClinicalMedicine
77 papers in training set
Top 0.9%
1.6%
17
Diabetologia
44 papers in training set
Top 0.5%
1.4%
18
Nutrients
67 papers in training set
Top 1%
1.2%
19
Journal of Racial and Ethnic Health Disparities
11 papers in training set
Top 0.2%
1.2%
20
European Journal of Public Health
21 papers in training set
Top 0.3%
1.2%
21
PNAS Nexus
159 papers in training set
Top 2%
1.2%
22
JAMA Network Open
130 papers in training set
Top 3%
1.0%
23
Journal of Affective Disorders
92 papers in training set
Top 1%
1.0%
24
PLOS Digital Health
106 papers in training set
Top 3%
1.0%
25
PLOS Medicine
110 papers in training set
Top 3%
0.9%
26
BMC Public Health
158 papers in training set
Top 5%
0.9%
27
Science Advances
1243 papers in training set
Top 29%
0.9%
28
Diabetes Research and Clinical Practice
11 papers in training set
Top 0.3%
0.9%
29
JAMIA Open
42 papers in training set
Top 1%
0.9%
30
The Lancet Regional Health - Americas
22 papers in training set
Top 0.5%
0.9%