Impact of control selection strategies on GWAS results: a study of prostate cancer in the UK Biobank
Lu, J.; Thygesen, J. H.; Beaumont, R. N.; Weedon, M. N.; Green, H.
Show abstract
As GWAS studies move from array-based genotyping to whole exome and genome sequencing, there is a significant increase in cost. Applying an appropriate technique for the selection of which controls to include, in large studies where more potential controls are available than needed for the study, may be a useful technique for minimising resource intensity while maintaining statistical power. We evaluated three control selection strategies in prostate cancer GWAS using 15,250 UK Biobank cases: (a) all controls, (b) matched controls, and (c) random selection. Both (b) and (c) achieved comparable power in detecting significant loci relative to (a), but matched controls (b) showed greater consistency in identifying leading SNPs. However, using (b) matched controls reduced discovery power by [~]30% compared with (a) all controls, highlighting a trade-off. Matching controls (1:4 ratio) offers a cost-effective approach for targeted SNP analysis across phenotypes but may miss novel associations. Availability and ImplementationR code for implementing matching and random control selection is provided and available on GitHub (https://github.com/Jingzhan-Lu/GWAS-Control-Selection).
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 92%
- Assessment of the Utility of Gene Positioning Biomarkers in the Stratification of Prostate Cancers 91%
- Impute.me: an open source, non-profit tool for using data from DTC genetic testing to calculate and interpret polygenic risk scores. 91%
Similar papers in this journal
- Validation and context-dependent effects of a prostate cancer polygenic risk score in the All of Us Research Program 92%
- A novel Bayesian fine-mapping model using a continuous global-local shrinkage prior with applications in prostate cancer analysis 92%
- Benchmarking Mendelian Randomization methods for causal inference using genome-wide association study summary statistics 91%
Similar papers in this journal
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 93%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 92%
- Investigating the Role of Neighborhood Socioeconomic Status and Germline Genetics on Prostate Cancer Risk 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.