Subtyping Social Determinants of Health in All of Us: Opportunities and Challenges in Integrating Multiple Datatypes for Precision Medicine
Bhavnani, S. K.; Zhang, W.; Bao, D.; Raji, M.; Kuo, Y.-f.; Schmidt, S.; Pappadis, M.; Bokov, A.; Reistetter, T.; Visweswaran, S.; Downer, B.
Show abstract
A.BackgroundSocial determinants of health (SDoH), such as financial resources and housing stability, account for between 30-55% of peoples health outcomes. While many studies have identified strong associations among specific SDoH and health outcomes, most people experience multiple SDoH that impact their daily lives. Analysis of this complexity requires the integration of personal, clinical, social, and environmental information from a large cohort of individuals that have been traditionally underrepresented in research, which is only recently being made available through the All of Us research program. However, little is known about the range and response of SDoH in All of Us, and how they co-occur to form subtypes, which are critical for designing targeted interventions. ObjectiveTo address two research questions: (1) What is the range and response to survey questions related to SDoH in the All of Us dataset? (2) How do SDoH co-occur to form subtypes, and what are their risk for adverse health outcomes? MethodsFor Question-1, an expert panel analyzed the range of SDoH questions across the surveys with respect to the 5 domains in Healthy People 2030 (HP-30), and analyzed their responses across the full All of Us data (n=372,397, V6). For Question-2, we used the following steps: (1) due to the missingness across the surveys, selected all participants with valid and complete SDoH data, and used inverse probability weighting to adjust their imbalance in demographics compared to the full data; (2) an expert panel grouped the SDoH questions into SDoH factors for enabling a more consistent granularity; (3) used bipartite modularity maximization to identify SDoH biclusters, their significance, and their replicability; (4) measured the association of each bicluster to three outcomes (depression, delayed medical care, emergency room visits in the last year) using multiple data types (surveys, electronic health records, and zip codes mapped to Medicaid expansion states); and (5) the expert panel inferred the subtype labels, potential mechanisms that precipitate adverse health outcomes, and interventions to prevent them. ResultsFor Question-1, we identified 110 SDoH questions across 4 surveys, which covered all 5 domains in HP-30. However, the results also revealed a large degree of missingness in survey responses (1.76%-84.56%), with later surveys having significantly fewer responses compared to earlier ones, and significant differences in race, ethnicity, and age of participants of those that completed the surveys with SDoH questions, compared to those in the full All of Us dataset. Furthermore, as the SDoH questions varied in granularity, they were categorized by an expert panel into 18 SDoH factors. For Question-2, the subtype analysis (n=12,913, d=18) identified 4 biclusters with significant biclusteredness (Q=0.13, random-Q=0.11, z=7.5, P<0.001), and significant replication (Real-RI=0.88, Random-RI=0.62, P<.001). Furthermore, there were statistically significant associations between specific subtypes and the outcomes, and with Medicaid expansion, each with meaningful interpretations and potential targeted interventions. For example, the subtype Socioeconomic Barriers included the SDoH factors not employed, food insecurity, housing insecurity, low income, low literacy, and low educational attainment, and had a significantly higher odds ratio (OR=4.2, CI=3.5-5.1, P-corr<.001) for depression, when compared to the subtype Sociocultural Barriers. Individuals that match this subtype profile could be screened early for depression and referred to social services for addressing combinations of SDoH such as housing insecurity and low income. Finally, the identified subtypes spanned one or more HP-30 domains revealing the difference between the current knowledge-based SDoH domains, and the data-driven subtypes. ConclusionsThe results revealed that the SDoH subtypes not only had statistically significant clustering and replicability, but also had significant associations with critical adverse health outcomes, which had translational implications for designing targeted SDoH interventions, decision-support systems to alert clinicians of potential risks, and for public policies. Furthermore, these SDoH subtypes spanned multiple SDoH domains defined by HP-30 revealing the complexity of SDoH in the real-world, and aligning with influential SDoH conceptual models such as by Dahlgren-Whitehead. However, the high-degree of missingness warrants repeating the analysis as the data becomes more complete. Consequently we designed our machine learning code to be generalizable and scalable, and made it available on the All of Us workbench, which can be used to periodically rerun the analysis as the dataset grows for analyzing subtypes related to SDoH, and beyond.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 94%
- A scoping review of fair machine learning techniques when using real-world data 94%
- Causal feature selection using a knowledge graph combining structured knowledge from the biomedical literature and ontologies: a use case studying depression as a risk factor for Alzheimer's disease 92%
Similar papers in this journal
Similar papers in this journal
- Characterization and Racial Stratification of Social Determinants of Health for Individuals with Type 2 Diabetes as Recorded in Electronic Health Records: Implications for Artificial Intelligence Development 94%
- Illustrating Potential Effects of Alternate Control Populations on Real-World Evidence-based Statistical Analyses 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 91%
Similar papers in this journal
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 93%
- Machine Learning Directed Interventions Associate with Decreased Hospitalization Rates in Hemodialysis Patients 91%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 90%
Similar papers in this journal
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 93%
- Cohort Profile: Genetic data in the German Socio-Economic Panel Innovation Sample (Gene-SOEP) 93%
- Impact of COVID-19 lockdown on psychosocial factors, health, and lifestyle in Scottish octogenarians: the Lothian Birth Cohort 1936 Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.