Back

Multi-PGS enhances polygenic prediction: weighting 937 polygenic scores

Albinana, C.; Zhu, Z.; Schork, A. J.; Ingason, A.; Aschard, H.; Brikell, I.; Bulik, C. M.; Pedersen, L. V.; Agerbo, E.; Grove, J.; Nordentoft, M.; Hougaard, D. M.; Werge, T.; Boerglum, A. D.; Mortensen, P. B.; McGrath, J. J.; Neale, B. M.; Prive, F.; Vilhjalmsson, B. J.

2022-09-17 genetic and genomic medicine
10.1101/2022.09.14.22279940 medRxiv
Show abstract

The predictive performance of polygenic scores (PGS) is largely dependent on the number of samples available to train the PGS. Increasing the sample size for a specific phenotype is expensive and takes time, but this sample size can be effectively increased by using genetically correlated phenotypes. We propose a framework to generate multi-PGS from thousands of publicly available genome-wide association studies (GWAS) with no need to individually select the most relevant ones. In this study, the multi-PGS framework increased prediction accuracy over single PGS for all included psychiatric disorders and other available outcomes, with prediction R2 increases of up to 9-fold for attention-deficit/hyperactivity disorder (ADHD) compared to a single PGS. We also generate multi-PGS for phenotypes without an existing GWAS and for case-case predictions, with up to 15-fold increases in prediction accuracy. We benchmark the multi-PGS framework against other methods and highlight its potential application to new emerging biobanks.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.