On PRS for complex polygenic trait prediction
Zhao, B.; Zou, F.
Show abstract
Polygenic risk score (PRS) is the state-of-art prediction method for complex traits using summary level data from discovery genome-wide association studies (GWAS). The PRS, as its name suggests, is designed for polygenic traits by aggregating small genetic effects from a large number of causal SNPs and thus is viewed as a powerful method for predicting complex polygenic traits by the genetics community. However, one concern is that the prediction accuracy of PRS in practice remains low with little clinical utility, even for highly heritable traits. Another practical concern is whether genome-wide SNPs should be used in constructing PRS or not. To address the two concerns, we investigate PRS both empirically and theoretically. We show how the performance of PRS is influenced by the triplet (n, p, m), where n, p, m are the sample size, the number of SNPs studied, and the number of true causal SNPs, respectively. For a given heritability, we find that i) when PRS is constructed with all p SNPs (referred as GWAS-PRS), its prediction accuracy is controlled by the p/n ratio; while ii) when PRS is built with a set of top-ranked SNPs that pass a pre-specified threshold (referred as threshold-PRS), its accuracy varies depending on how sparse the true genetic signals are. Only when m is magnitude smaller than n, or genetic signals are sparse, can threshold-PRS perform well and outperform GWAS-PRS. Our results demystify the low performance of PRS in predicting highly polygenic traits, which will greatly increase researchers aware-ness of the power and limitations of PRS, and clear up some confusion on the clinical application of PRS.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bayesian Hierarchical Hypothesis Testing in Large-Scale Genome-Wide Association Analysis 96%
- Reflection Knockoffs via Householder Reflection: Applications in Proteomics and Genetic Fine Mapping 96%
- Using encrypted genotypes and phenotypes for collaborative genomic analyses to maintain data confidentiality 95%
Similar papers in this journal
- Identifying causal genotype-phenotype relationships for population-sampled parent-child trios 97%
- A Two-Sample Robust Bayesian Mendelian Randomization Method Accounting for Linkage Disequilibrium and Idiosyncratic Pleiotropy with Applications to the COVID-19 Outcome 96%
- Improved two-step testing of genome-wide gene-environment interactions 96%
Similar papers in this journal
- Estimating the overall fraction of phenotypic variance attributed to high-dimensional predictors measured with error 95%
- The winner's curse under dependence: repairing empirical Bayes using convoluted densities 95%
- A mixed-model approach for powerful testing of genetic associations with cancer risk incorporating tumor characteristics 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.