Geographically Weighted Machine Learning for Spatial Prediction of Cancer Prevalence in the United States: A Mixed Method Approach
Sadeghi Naieni Fard, F.; Oppong, J. R.; Tiwari, C.; Boakye, K.; Fard, F.
Show abstract
Cancer prevalence is distributed unevenly across regions and caused by the interaction of multiple risk factors. Previous studies focused on the use of global modeling techniques to predict cancer at the county level that overlooks important spatial differences. This study aims to develop geographically weighted machine learning models to predict cancer prevalence at the census tract level in the United States and identify local determinants of cancer burden. First, a scoping review was conducted to find a list of measurable drivers of cancer in the United States. Using this list, the data of these variables for 84415 census tracts were obtained from the Center for Disease Control and Prevention PLACES dataset and other publicly accessible resources. Then, several predictive models, including Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR), as well as Random Forest, XGBoost, and Deep Neural Network and their geographically weighted counterparts, were developed and compared using the Coefficient of Determination, Root Mean Square Error, and Absolute Error. Results presented that geographically weighted models outperformed other methods, and geographically weighted XGBoost achieved the strongest and most consistent overall performance with pseudo-R2 ranging between 0.89 and 0.98. Feature importance analysis of this model illustrated that most important cancer drivers changed location by location. Aged people, racial composition, preventative behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol were determined as influential predictors, although their relative importance varied across regions. These findings revealed the value of localized models at a small geographic scale to identify regional cancer risk patterns and help the allocation of proper resources to hotspot areas. Keywords: Cancer prevalence, Census tracts, geographically weighted machine learning models, Deep neural network, XGBoost, Random Forest, Ordinary Least Squares, risk factor, determinant
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 92%
- Geographic Disparities and Determinants of COVID-19 Incidence Risk in the Greater St. Louis Area, Missouri 92%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 92%
Similar papers in this journal
- Isolation Considered Epidemiological Model for the Prediction of COVID-19 Trend in Tokyo, Japan 92%
- A Multivariate Forecasting Model for the COVID-19 Hospital Census Based on Local Infection Incidence 91%
- Estimating COVID-19 Hospitalizations in the United States with surveillance data using a Bayesian Hierarchical model 89%
Similar papers in this journal
- Machine Learning Based Clinical Decision Support System for Early COVID-19 Mortality Prediction 91%
- Unequal Benefits: The Effects of Health Insurance Integration on Consumption Inequality in Rural China 91%
- Nowcasting and Forecasting the Spread of COVID-19 and Healthcare Demand In Turkey, A Modelling Study 91%
Similar papers in this journal
- A Comprehensive County Level Framework to Identify Factors Affecting Hospital Capacity and Predict Future Hospital Demand 92%
- A deep learning approach for Pan-Renal Cell Carcinoma classification and survival prediction from histopathology images 92%
- A multipurpose machine learning approach to predict COVID-19 negative prognosis in Sao Paulo, Brazil 92%
Similar papers in this journal
- Evolving Patterns of COVID-19 Mortality in US Counties: A Longitudinal Study of Healthcare, Socioeconomic, and Vaccination Associations 92%
- Measuring the impact of nonpharmaceutical interventions on the SARS-CoV-2 pandemic at a city level: An agent-based computational modeling study of the City of Natal 90%
- Changes in the prevalence of the common risk factors for non-communicable diseases in Uganda between 2014 and 2023: Informed by nationally representative cross-sectional surveys 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.