A Combined Predictive and Causal Approach for Neighborhood-Level Diabetes Detection
Noaeen, M.; Rostami, A.; Ghanem, I.; Saarela, O.; Keshavjee, K.; Brook, J. R.; Shakeri, Z.
Show abstract
ObjectiveDevelop a neighborhood-level framework using machine learning and causal inference to identify socioeconomic and behavioral drivers of Type 2 diabetes for targeted public health interventions. Materials and MethodsData from 1,149 Census Tracts in Toronto were integrated, linking demographic, health, and marginalization indices. Seven machine learning models classified neighborhoods with high diabetes prevalence. Feature engineering mitigated skewness and correlation, while Causal Forests estimated the Conditional Average Treatment Effect (CATE,{tau} ) for predictors such as work stress, smoking, and mental health. ResultsPredictive models achieved over 90% recall and high AUC metrics on both test and external validation datasets. Key predictors included obesity, overweight status, physical activity, and log-transformed median age. Causal analysis further indicated that elevated work stress ({tau} = 0.312) and daily smoking ({tau} = 0.155) increased diabetes risk, while stronger mental health ({tau} {approx} -1.1) was protective. DiscussionWhile genetic and clinical factors often dominate the conversation on diabetes, data is often restricted to confirmed diagnoses or not readily available for prevalence analyses. Our study shows how neighborhood contexts, including walkability, stress levels, and socioeconomic differences, help drive rising disease rates. We integrated machine learning classifiers with causal inference to examine how interventions, such as active transportation and adjusted work stress, could shift diabetes risk. ConclusionThis integrated method offers a blueprint for precision public health by clarifying how modifiable neighborhood factors affect diabetes risk. It can help tailor interventions to community needs and is applicable to other areas facing similar chronic disease challenges.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adherence and sustainability of interventions informing optimal control against COVID-19 pandemic 89%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 89%
- Systematic review of precision subclassification of type 2 diabetes 88%
Similar papers in this journal
- The impact of social and environmental extremes on cholera time varying reproduction number in Nigeria 90%
- Vaccine Rollout Strategies: The Case for Vaccinating Essential Workers Early 89%
- Mobility changes following COVID-19 stay-at-home policies varied by socioeconomic measures: An observational study in Ontario, Canada 89%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Identifying communities at risk for COVID-19-related burden across 500 U.S. Cities and within New York City 93%
- Uncovering clinical risk factors and prediction of severe COVID-19: A machine learning approach based on UK Biobank data 91%
- Subphenotyping of COVID-19 patients at pre-admission towards anticipated severity stratification: an analysis of 778 692 Mexican patients through an age-gender unbiased meta-clustering technique 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.