Back

Comparing different methods of estimating GWAS heritability with a new approach using only summary statistics

Salehi, E.

2023-10-03 genetics
10.1101/2023.10.02.560406 bioRxiv
Show abstract

So far SNP heritability ([Formula];variance explained by all SNP s used in genome-wide association study) has explained most of genetic variation for many traits but still there is a gap between GWAS heritability ([Formula]; variance explained by genome-wide significant SNPs) and [Formula] that is named hidden heritability. There are several methods for estimating [Formula] (linear_mixed_model (LMM), PRS, multiple_linear_regression (MLR) and simple_linear_regression(SLR)). However, it is unclear which methods are more accurate under different circumstances. This study proposes a PRS based method for estimating [Formula] that uses pseudo summary statistics. It compares this method with existing methods using both simulated and real data (10 traits from UKBB) to determine when they are realistic and can be trusted as a final estimate. Simulation results showed that PRS-based methods underestimate [Formula] near 20% when considering all causal SNPs. But they are relatively accurate when using a subset of causal SNPs. Their performance is much better than SLR method for all 10 traits, although when applied to real data, they do not follow a stable trend of overestimation or underestimation compared to the base model (LMM). My suggestion is to use LMM or adjusted_R2 from MLR for reporting [Formula] when an independent data set is available. In cases where only summary statistics is available, the PRS-PSS is relatively an accurate alternative, especially compared to SLR, which tends to overestimate [Formula] by 20-50% when applying it on real data.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.