Back

Determining Personalized Community Health Needs by Feature Selection and Clustering

Agar-Johnson, M.

2020-02-23 health economics
10.1101/2020.02.21.20024612 medRxiv
Show abstract

The Center for Disease Control, through the Community Health Data Initiative (CHDI), has released a large dataset by county detailing the overall health indicators, demographics, and major risk factors and causes of morbidity and mortality in the US. In order to address the heterogeneity of community healthcare in the US, k-Means clustering was performed on the CHDI dataset to determine community subtypes in terms of health challenges and outcomes. The optimal number of eight clusters was determined by the Elbow Method, and clusters were analyzed to determine significant differences in demographic. In order to determine community-specific healthcare solutions and directions, feature selection and modeling of healthcare outcomes was performed for each of the eight subtypes using LASSO regression. It was determined that different features significantly impact health outcomes in the different clusters, providing information about the unique health challenges faced by these different types of communities. LASSO regression using the entire unclustered dataset yielded significantly poorer results on the sub-clusters in terms of model performance, further supporting the claim that modeling community-specific needs is a vital step for delivering accurate and adequate community healthcare. These results have the potential to inform policymaking at the local/municipal level, as well as inform the approaches taken by primary practitioners to address community needs.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.