Explainable machine learning models to understand determinants of COVID-19 mortality in the United States
Mathur, P.; Sethi, T.; Mathur, A.; Khanna, A. K.; Maheshwari, K.; Cywinski, J. B.; Dua, S.; Papay, F.
Show abstract
BackgroundCOVID-19 is now one of the leading causes of mortality amongst adults in the United States for the year 2020. Multiple epidemiological models have been built, often based on limited data, to understand the spread and impact of the pandemic. However, many geographic and local factors may have played an important role in higher morbidity and mortality in certain populations. ObjectiveThe goal of this study was to develop machine learning models to understand the relative association of socioeconomic, demographic, travel, and health care characteristics of different states across the United States and COVID-19 mortality. MethodsUsing multiple public data sets, 24 variables linked to COVID-19 disease were chosen to build the models. Two independent machine learning models using CatBoost regression and random forest were developed. SHAP feature importance and a Boruta algorithm were used to elucidate the relative importance of features on COVID-19 mortality in the United States. ResultsFeature importances from both the categorical models, i.e., CatBoost and random forest consistently showed that a high population density, number of nursing homes, number of nursing home beds and foreign travel were strongest predictors of COVID-19 mortality. Percentage of African American amongst the population was also found to be of high importance in prediction of COVID-19 mortality whereas racial majority (primarily, Caucasian) was not. Both models fitted the data well with a training R2 of 0.99 and 0.88 respectively. The effect of median age,median income, climate and disease mitigation measures on COVID-19 related mortality remained unclear. ConclusionsCOVID-19 policy making will need to take population density, pre-existing medical care and state travel policies into account. Our models identified and quantified the relative importance of each of these for mortality predictions using machine learning.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Application of Elastic Net Regression for Modeling COVID-19 Sociodemographic Risk Factors 93%
- Understanding mental health trends during COVID-19 pandemic in the United States using network analysis 93%
- Geographic Disparities and Determinants of COVID-19 Incidence Risk in the Greater St. Louis Area, Missouri 92%
Similar papers in this journal
- Predicting mortality in SARS-COV-2 (COVID-19) positive patients in the inpatient setting using a Novel Deep Neural Network 91%
- Predictive Model with Analysis of the Initial Spread of COVID-19 in India 90%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 90%
Similar papers in this journal
- A Comprehensive County Level Framework to Identify Factors Affecting Hospital Capacity and Predict Future Hospital Demand 94%
- Predicting Car Accident Severity in Northwest Ethiopia: A Machine Learning Approach Leveraging Driver, Environmental, and Road Conditions 92%
- A multipurpose machine learning approach to predict COVID-19 negative prognosis in Sao Paulo, Brazil 92%
Similar papers in this journal
- A Multivariate Forecasting Model for the COVID-19 Hospital Census Based on Local Infection Incidence 92%
- Estimating COVID-19 Hospitalizations in the United States with surveillance data using a Bayesian Hierarchical model 91%
- Isolation Considered Epidemiological Model for the Prediction of COVID-19 Trend in Tokyo, Japan 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.