Back

Estimating Disorder Probability Based on Polygenic Prediction Using the BPC Approach

Uffelmann, E.; Major Depressive Disorder Working Group of the Psychiatric Genomics Consortium, ; Schizophrenia Working Group of the Psychiatric Genomics Consortium, ; Price, A. L.; Posthuma, D.; Peyrot, W. J.

2024-01-13 genetic and genomic medicine
10.1101/2024.01.12.24301157 medRxiv
Show abstract

Polygenic Scores (PGSs) summarize an individuals genetic propensity for a given trait in a single value based on SNP effect sizes derived from Genome-Wide Association Study (GWAS) results. Methods have been developed that apply Bayesian approaches to improve the prediction accuracy of PGSs through optimization of estimated effect sizes. While these methods are generally well-calibrated for continuous traits (implying the predicted values are, on average, equal to the true trait values), they are not well-calibrated for binary disorder traits in ascertained samles. This is a problem because well-calibrated PGSs are needed to reliably compute the absolute disorder probability for an individual to facilitate future clinical implementation. Here, we introduce the Bayesian polygenic score Probability Conversion (BPC) approach, which computes an individuals predicted disorder probability using GWAS summary statistics, an existing Bayesian PGS method (e.g., PRScs, SBayesR), the individuals genotype data, and a prior disorder probability (which can be specified flexibly, based on e.g., literature, small reference samples, or prior elicitation). The BPC approach transforms the PGS to its underlying liability scale, computes the variances of the PGS in cases and controls, and applies Bayes Theorem to compute the absolute disorder probability; it is practical in its application as it does not require a tuning sample with both genotype and phenotype data. We applied the BPC approach to extensive simulated data and empirical data of nine disorders. The BPC approach yielded well-calibrated results that were consistently better than the results of another recently published approach.

Published in Nature Communications (predicted rank #5) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.