Back

Using Obviously-Related Instrumental Variables to Increase the Predictive Power of Polygenic Scores

van Kippersluis, H.; Biroli, P.; Galama, T. J.; von Hinke, S.; Meddens, S. F. W.; Muslimova, D.; Pereira, R.; Rietveld, C. A.

2021-08-03 genetics
10.1101/2021.04.09.439157 bioRxiv
Show abstract

Measurement error in polygenic indices (PGIs) attenuates the estimation of their effects in regression models. While this measurement error shrinks with growing Genome-wide Association Study (GWAS) sample sizes, the marginal returns to bigger sample sizes are rapidly decreasing. We analyze and compare two alternative approaches to reduce measurement error: Obviously Related Instrumental Variables (ORIV) and the PGI Repository Correction (PGI-RC). Through simulations, we show that both approaches outperform the typical (meta-analysis based) PGI in terms of bias and root mean squared error. Between families, the PGI-RC performs slightly better than ORIV, unless the prediction sample is very small (N < 1, 000), or when there is considerable assortative mating. Within families, ORIV is the default choice since the PGI-RC is not available in this setting. We verify the empirical validity of the simulations by predicting educational attainment (EA) and height in a sample of siblings from the UK Biobank. We show that applying ORIV between families increases the standardized effect of the PGI by 12% (height) and by 22% (EA) compared to a meta-analysis-based PGI, yet remains slightly below the PGI-RC estimates. Furthermore, within-family ORIV regression provides the tightest lower bound for the direct genetic effect, increasing the lower bound for the direct genetic effect on EA from 0.14 to 0.18, and for height from 0.54 to 0.61 compared to a meta-analysis-based PGI.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.