Selecting Covariates for Genome-Wide Association Studies
Dor, E.; Margaliot, I.; Brandes, N.; Zuk, O.; Linial, M.; Rappoport, N.
Show abstract
The choice of which covariates to include in a Genome-Wide Association Study (GWAS) is important since it affects the ability to detect true association signal of variants, to correct for confounders and avoid false positives, and the running time of the analysis. Commonly used covariates include age, sex, genotyping batches, genotyping array type, as well as an arbitrary number of Principal Components (PCs) used to adjust for population structure. Despite the importance of this issue, there is no consensus or clear guidelines for the right choice of covariates. Therefore, studies typically employ heuristics for their choice with no clear justification. Here, we explore the dependence of the GWAS analysis results on the choice of covariates for a wide range of quantitative and binary human phenotypes. We propose guidelines for covariates choice based on the phenotypes type (quantitative vs. disease), the heritability, and the disease prevalence, with the goal of maximizing the statistical power to detect true associations and fit accurate polygenic scores while avoiding spurious associations and minimizing computation time. We analyze 36 traits in the UK-Biobank dataset. We show that the genotype batch and assessment center can be safely removed as covariates, thus significantly reducing the GWAS computational burden for these traits.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 96%
- Noise-augmented directional clustering of genetic association data identifies distinct mechanisms underlying obesity 95%
Similar papers in this journal
- Genome-wide haplotype association study in imaging genetics using whole-brain sulcal openings of 16,304 UK Biobank subjects. 95%
- A Tool for Translating Polygenic Scores onto the Absolute Scale Using Summary Statistics 93%
- Exploiting Family History in Aggregation Unit-based Genetic Association Tests 93%
Similar papers in this journal
- Assumptions about frequency-dependent architectures of complex traits bias measures of functional enrichment 95%
- Benchmarking statistical methods for analyzing parent-child dyads in genetic association studies 94%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.