Using summary data to detect and quantify ascertainment in biobanks
Olasege, B. S.; Campos, A. I.; Sidorenko, J.; Lin, T.; Barry, C.-J. S.; Maseras, G. T.; Vilhjalmsson, B. J.; Wray, N. R.; Hivert, V.; Yengo, L.
Show abstract
Non-random participation in genetic studies can bias associations between genetic variants and outcomes. Existing methods to detect ascertainment bias often require individual-level data, thus limiting their broad applicability. Here, we introduce a summary-statistics-based method to detect and quantify ascertainment bias in large-scale genetic studies. Our method estimates a parameter,{theta} , which captures deviations in the mean polygenic score (PGS) of an ascertained sample relative to its expectation across non-ascertained or differentially ascertained references. We show through extensive simulations that our method is robust to population stratification and reference misspecification unlike naive mean PGS comparison. When applied to 21 traits across 11 large-scale biobanks, our method recapitulates known patterns of ascertainment and detects new evidence of ascertainment on genetic susceptibility to depression, height and blood pressure in many biobanks. Overall, our framework enables systematic assessment of ascertainment directly from summary statistics and provides a scalable tool for evaluating representativeness in large scale genetic studies.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimating disease heritability from complex pedigrees allowing for ascertainment and covariates 96%
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 95%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 95%
Similar papers in this journal
- The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup 96%
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 95%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 95%
Similar papers in this journal
- MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data 96%
- Leveraging a machine learning derived surrogate phenotype to improve power for genome-wide association studies of partially missing phenotypes in population biobanks 95%
- Reconciling S-LDSC and LDAK functional enrichment estimates 95%
Similar papers in this journal
- A Comprehensive Evaluation of Methods for Mendelian Randomization Using Realistic Simulations and an Analysis of 38 Biomarkers for Risk of Type-2 Diabetes 94%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 94%
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.