Reducing Information and Selection Bias in EHR-Linked Biobanks via Genetics-Informed Multiple Imputation and Sample Weighting
Salvatore, M.; Kundu, R.; Du, J.; Friese, C. R.; Mondul, A. M.; Hanauer, D. A.; Lu, H.; Pearce, C. L.; Mukherjee, B.
Show abstract
Electronic health records (EHRs) are valuable for public health and clinical research but are prone to many sources of bias, including missing data and non-probability selection. Missing data in EHRs is complex due to potential non-recording, fragmentation, or clinically informative absences. This study explores whether polygenic risk score (PRS)-informed multiple imputation for missing traits, combined with sample weighting, can mitigate missing data and selection biases in estimating disease-exposure associations. Simulations were conducted for missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR) conditions under different sampling mechanisms. PRS-informed multiple imputation showed generally lower bias, particularly when combined with sample weighting. For example, in biased samples of 10,000 with exposure and outcome MAR data, PRS-informed imputation had lower percent bias (3.8%) and better coverage rate (0.883) compared to PRS-uninformed (4.5%; 0.877) and complete case analyses (10.3%; 0.784) in covariate-adjusted, weighted, multiple imputation scenarios. In a case study using Michigan Genomics Initiative (n=50,026) data, PRS-informed imputation aligned more closely with a sample-weighted All of Us-derived benchmark than analyses ignoring missing data and selection bias. Researchers should consider leveraging genetic data and sample weighting to address biases from missing data and non-probability sampling in biobanks.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Within-family studies for Mendelian randomization: avoiding dynastic, assortative mating, and population stratification biases 95%
- A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk 93%
- Comprehensive genomic analysis of dietary habits in UK Biobank identifies hundreds of genetic loci and establishes causal relationships between educational attainment and healthy eating 93%
Similar papers in this journal
- A Hierarchical Approach Using Marginal Summary Statistics for Multiple Intermediates in a Mendelian Randomization or Transcriptome Analysis 91%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 91%
- A structural mean modelling Mendelian randomization approach to investigate the lifecourse effect of adiposity: applied and methodological considerations 91%
Similar papers in this journal
- Epigenome-wide association study of incident type 2 diabetes in Black and White participants from the Atherosclerosis Risk in Communities Study 95%
- Phenotype-based targeted treatment of SGLT2 inhibitors and GLP-1 receptor agonists in type 2 diabetes 94%
- The power of TOPMed imputation for the discovery of Latino enriched rare variants associated with type 2 diabetes 92%
Similar papers in this journal
- Risk factors affecting polygenic score performance across diverse cohorts 93%
- Nuclear magnetic resonance-based metabolomics with machine learning for predicting progression from prediabetes to diabetes 93%
- Serum proteomic profiling of physical activity reveals CD300LG as a novel exerkine with a potential causal link to glucose homeostasis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.