Population stratification in GWAS meta-analysis should be standardized to the best available reference datasets
Sarmanova, A.; Morris, T. T.; Lawson, D. J.
Show abstract
Population stratification has recently been demonstrated to bias genetic studies even in relatively homogeneous populations such as within the British Isles. A key component to correcting for stratification in genome-wide association studies (GWAS) is accurately identifying and controlling for the underlying structure present in the sample. Meta-analysis across cohorts is increasingly important for achieving very large sample sizes, but comes with the major disadvantage that each individual cohort corrects for different population stratification. Here we demonstrate that correcting for structure against an external reference adds significant value to meta-analysis. We treat the UK Biobank as a collection of smaller studies, each of which is geographically localised. We provide software to standardize an external dataset against a reference, provide the UK Biobank principal component loadings for this purpose, and demonstrate the value of this with an analysis of the geographically sampled ALSPAC cohort.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-resolution portability of 245 polygenic scores when derived and applied in the same cohort 97%
- Evaluating Multi-Ancestry Genome-Wide Association Methods: Statistical Power, Population Structure, and Practical Implications 96%
- The individual and global impact of copy number variants on complex human traits 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Estimating effects of parents' cognitive and non-cognitive skills on offspring education using polygenic scores 96%
- A novel method for an unbiased estimate of cross-ancestry genetic correlation using individual-level data 96%
- Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.