Explainable machine learning for health disparities: type 2 diabetes in the All of Us research program
Kambara, M. S.; Chukka, O.; Choi, K. J.; Tsenum, J.; Gupta, S.; English, N. J.; Jordan, I. K.; Marino-Ramirez, L.
Show abstract
Type 2 diabetes (T2D) is a disease with high morbidity and mortality and a disproportionate impact on minority groups. Machine learning (ML) is increasingly used to characterize T2D risk factors; however, it has not been used to study T2D health disparities. Our objective was to use explainable ML methods to discover and characterize T2D health disparity risk factors. We applied SHapley Additive exPlanations (SHAP), a new class of explainable ML methods that provide interpretability to ML classifiers, to this end. ML classifiers were used to model T2D risk within and between self-identified race and ethnicity (SIRE) groups, and SHAP values were calculated to quantify the effect of T2D risk factors. We then stratified SHAP values by SIRE to quantify the effect of T2D risk factors on prevalence differences between groups. We found that ML classifiers (random forest, lightGBM, and XGBoost) accurately modeled T2D risk and recaptured the observed prevalence differences between SIRE groups. SHAP analysis showed the top seven most important T2D risk factors for all SIRE groups were the same, with the order of importance for features differing between groups. SHAP values stratified by SIRE showed that income, waist circumference, and education best explain the higher prevalence of T2D in the Black or African American group, compared to the White group, whereas income, education and triglycerides best explain the higher prevalence of T2D in the Hispanic or Latino group. This study demonstrates that explainable ML can be used to elucidate health disparity risk factors and quantify their group-specific effects. Author SummaryWhile machine learning (ML) methods hold great promise for epidemiological studies, their practical utility is limited by interpretability. Increasingly complex ML models are great at predicting disease risk, but how they arrive at a given prediction is often obscured by model complexity. Explainable ML is an emerging discipline that seeks to render ML models more transparent by elucidating how and why input features contribute to output predictions. This study reports a novel application of explainable ML to epidemiology, focusing on type 2 diabetes (T2D) as a paradigm of health disparities. We found that ML classifiers were able to accurately model T2D disparities, for a large cohort of Black, Hispanic, and White Americans, and explainable ML revealed which risk factors contributed to the observed disparities and how. The results demonstrate that explainable ML can be a powerful tool for the discovery and characterization of health disparity risk factors.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Actionable absolute risk prediction of atherosclerotic cardiovascular disease: a behavior-management approach based on data from 464,547 UK Biobank participants 93%
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 92%
- Development and Validation of the Michigan Chronic Disease Simulation Model (MICROSIM) 91%
Similar papers in this journal
- Age at Menarche and Coronary Artery Disease Risk: Divergent Associations with Different Sources of Variation 92%
- Circulating metabolic biomarkers are consistently associated with incident type 2 diabetes in Asian and European populations – a metabolomics analysis in five prospective cohorts 92%
- Prediabetes as a risk factor for all-cause and cause-specific mortality: a prospective analysis of 115,919 adults without diabetes in Mexico City 92%
Similar papers in this journal
- Nationwide prediction of type 2 diabetes comorbidities 93%
- What multiple Mendelian randomization approaches reveal about obesity and gout 93%
- Can machine learning improve risk prediction of incident hypertension? An internal method comparison and external validation of the Framingham risk model using HUNT Study data 92%
Similar papers in this journal
- Dietaryindex: A User-Friendly and Versatile R Package for Standardizing Dietary Pattern Analysis in Epidemiological and Clinical Studies 90%
- Adjustment for energy intake in nutritional research: a causal inference perspective 90%
- Monitoring Body Composition Change for Intervention Studies with Advancing 3D Optical Imaging Technology in Comparison to Dual-Energy X-Ray Absorptiometry 89%
Similar papers in this journal
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 91%
- Contrasting factors associated with COVID-19-related ICU admission and death outcomes in hospitalised patients by means of Shapley values 91%
- Cross-fitted instrument: a blueprint for one-sample Mendelian Randomization 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.