Pitfalls in estimating and interpreting the contribution of ultra-rare genetic variants to the heritability of complex traits
Wang, H.; Wainschtein, P.; Sidorenko, J.; Fikere, M.; Zhang, Y.; Kemper, K. E.; Zheng, Z.; Hivert, V.; Zeng, J.; Goddard, M. E.; Visscher, P. M.; Yengo, L.
Show abstract
Assessing the contribution of ultra-rare variants (minor allele frequency <0.01%) to the heritability of complex traits remains challenging due to limited understanding of potential biases. Here, we focus on singletons (that is, variants observed only once in the study sample), the most abundant class of ultra-rare variants, to showcase various confounders of heritability estimates and underline pitfalls in their interpretation. We show through theory, simulations, and analysis of 5,330,210 exome-sequenced singletons in 305,813 unrelated European-ancestry individuals in the UK Biobank that (i) population stratification induces both upward and downward biases in singleton-based heritability estimates (), (ii) estimates capture non-additive genetic effects, and (iii) asymptotic standard errors of estimates from likelihood-based procedures are generally mis-calibrated when traits are not normally distributed. We further showcase these biases in real-data analyses of 22 quantitative phenotypes and report, after accounting for these pitfalls, significant estimate for number of children (3.4%), peak expiratory flow (1.9%), red blood cell count (2.5%), white blood cell count (1.9%) and heel bone mineral density (2.4%). Overall, our study provides recommendations for robust inference of heritability from ultra rare variants and underscores that reliable estimates for ordinal and binary traits will require far larger sample sizes and improved methods, given that confounding in these traits remains difficult to detect and correct
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimation of non-additive genetic variance in human complex traits from a large sample of unrelated individuals 98%
- Estimating disease heritability from complex pedigrees allowing for ascertainment and covariates 97%
- Evaluating Multi-Ancestry Genome-Wide Association Methods: Statistical Power, Population Structure, and Practical Implications 97%
Similar papers in this journal
- A novel method for an unbiased estimate of cross-ancestry genetic correlation using individual-level data 97%
- Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations 97%
- Improved analyses of GWAS summary statistics by reducing data heterogeneity and errors 97%
Similar papers in this journal
Similar papers in this journal
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 97%
- Calibrated prediction intervals for polygenic scores across diverse contexts 96%
- Combining case-control status and family history of disease increases association power 96%
Similar papers in this journal
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 95%
- Inclusion of Variants Discovered from Diverse Populations Improves Polygenic Risk Score Transferability 95%
- Polygenic risk score prediction accuracy convergence 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.