Publicly Available Privacy-preserving Benchmarks for Polygenic Prediction
Witteveen, M. J.; Pedersen, E. M.; Meijsen, J.; Andersen, M. R.; Prive, F.; Speed, D.; Vilhjalmsson, B. J.
Show abstract
Recently, several new approaches for creating polygenic scores (PGS) have been developed and this trend shows no sign of abating. However, it has thus far been challenging to determine which approaches are superior, as different studies report seemingly conflicting benchmark results. This heterogeneity in benchmark results is in part due to different outcomes being used, but also due to differences in the genetic variants being used, data preprocessing, and other quality control steps. As a solution, a publicly available benchmark for polygenic prediction is presented here, which allows researchers to both train and test polygenic prediction methods using only summary-level information, thus preserving privacy. Using simulations and real data, we show that model performance can be estimated with accuracy, using only linkage disequilibrium (LD) information and genome-wide association summary statistics for target outcomes. Finally, we make this PGS benchmark - consisting of 8 outcomes, including somatic and psychiatric disorders - publicly available for researchers to download on our PGS benchmark platform (http://www.pgsbenchmark.org). We believe this benchmark can help establish a clear and unbiased standard for future polygenic score methods to compare against.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 95%
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 95%
- Scalable probabilistic PCA for large-scale genetic variation data 95%
Similar papers in this journal
- Fast Kernel-based Association Testing of non-linear genetic effects for Biobank-scale data 95%
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 95%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.