Fine-tuning Polygenic Risk Scores with GWAS Summary Statistics
Zhao, Z.; Yi, Y.; Wu, Y.; Zhong, X.; Lin, Y.; Hohman, T. J.; Fletcher, J.; Lu, Q.
Show abstract
Polygenic risk scores (PRSs) have wide applications in human genetics research. Notably, most PRS models include tuning parameters which improve predictive performance when properly selected. However, existing model-tuning methods require individual-level genetic data as the training dataset or as a validation dataset independent from both training and testing samples. These data rarely exist in practice, creating a significant gap between PRS methodology and applications. Here, we introduce PUMAS (Parameter-tuning Using Marginal Association Statistics), a novel method to fine-tune PRS models using summary statistics from genome-wide association studies (GWASs). Through extensive simulations, external validations, and analysis of 65 traits, we demonstrate that PUMAS can perform a variety of model-tuning procedures (e.g. cross-validation) using GWAS summary statistics and can effectively benchmark and optimize PRS models under diverse genetic architecture. On average, PUMAS improves the predictive R2 by 205.6% and 62.5% compared to PRSs with arbitrary p-value cutoffs of 0.01 and 1, respectively. Applied to 211 neuroimaging traits and Alzheimers disease, we show that fine-tuned PRSs will significantly improve statistical power in downstream association analysis. We believe our method resolves a fundamental problem without a current solution and will greatly benefit genetic prediction applications.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bayesian transcriptome-wide association study method leveraging both cis- and trans- eQTL information through summary statistics 97%
- Non-parametric polygenic risk prediction using partitioned GWAS summary statistics 96%
- A method to map and interpret pleiotropic loci using summary statistics of multiple traits 96%
Similar papers in this journal
- Scalable Bayesian functional GWAS method accounting for multivariate quantitative functional annotations with applications to studying Alzheimer’s disease 97%
- Polygenic risk score prediction accuracy convergence 95%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 94%
Similar papers in this journal
- Identity-by-descent mapping using multi-individual IBD with genome-wide multiple testing adjustment 95%
- Meta-MultiSKAT: Multiple phenotype meta-analysis for region-based association test 93%
- RetroFun-RVS: a retrospective family-based framework for rare-variant analysis incorporating functional annotations 93%
Similar papers in this journal
- Identification of putative causal loci in whole-genome sequencing data via knockoff statistics 97%
- Co-expression-wide association studies link genetically regulated interactions with complex traits 96%
- Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics 95%
Similar papers in this journal
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 95%
- Summary statistics from large-scale gene-environment interaction studies for re-analysis and meta-analysis 95%
- MR Corge: Sensitivity analysis of Mendelian randomization based on the core gene hypothesis for polygenic exposures 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.