To weight or not to weight? Studying the effect of selection bias in three large EHR-linked biobanks
Salvatore, M.; Kundu, R.; Shi, X.; Friese, C. R.; Lee, S.; Fritsche, L. G.; Mondul, A. M.; Hanauer, D. A.; Pearce, C. L.; Mukherjee, B.
Show abstract
ObjectiveTo explore the role of selection bias adjustment by weighting electronic health record (EHR)-linked biobank data for commonly performed analyses. Materials and methodsWe mapped diagnosis (ICD code) data to standardized phecodes from three EHR-linked biobanks with varying recruitment strategies: All of Us (AOU; n=244,071), Michigan Genomics Initiative (MGI; n=81,243), and UK Biobank (UKB; n=401,167). Using 2019 National Health Interview Survey data, we constructed selection weights for AOU and MGI to be more representative of the US adult population. We used weights previously developed for UKB to represent the UKB-eligible population. We conducted four common descriptive and analytic tasks comparing unweighted and weighted results. ResultsFor AOU and MGI, estimated phecode prevalences decreased after weighting (weighted-unweighted median phecode prevalence ratio [MPR]: 0.82 and 0.61), while UKBs estimates increased (MPR: 1.06). Weighting minimally impacted latent phenome dimensionality estimation. Comparing weighted versus unweighted PheWAS for colorectal cancer, the strongest associations remained unaltered and there was large overlap in significant hits. Weighting affected the estimated log-odds ratio for sex and colorectal cancer to align more closely with national registry-based estimates. DiscussionWeighting had limited impact on dimensionality estimation and large-scale hypothesis testing but impacted prevalence and association estimation more. Results from untargeted association analyses should be followed by weighted analysis when effect size estimation is of interest for specific signals. ConclusionEHR-linked biobanks should report recruitment and selection mechanisms and provide selection weights with defined target populations. Researchers should consider their intended estimands, specify source and target populations, and weight EHR-linked biobank analyses accordingly.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk 96%
- Medical history predicts phenome-wide disease onset 95%
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 94%
Similar papers in this journal
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 94%
- Exome-by-phenome-wide rare variant gene burden association with electronic health record phenotypes 94%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 94%
Similar papers in this journal
Similar papers in this journal
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 93%
- Association between household composition and severe COVID-19 outcomes in older people by ethnicity: an observational cohort study using the OpenSAFELY platform 92%
- Genetic evidence for causal relationships between age at natural menopause and the risk of aging-associated adverse health outcomes 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.