Harnessing population-specific protein truncating variants to improve the annotation of loss-of-function alleles
Skitchenko, R. K.; Kornienko, J. S.; Maksiutenko, E. M.; Glotov, A. S.; Predeus, A. V.; Barbitoff, Y. A.
Show abstract
Accurate annotation of putative loss-of-function (pLoF) variants is an important problem in human genomics and disease, which recently drew substantial attention. Since such variants in disease-related genes are under strong negative selection, their frequency across major ancestral groups is expected to be highly similar. In this study, we tested this assumption by systematically assessing the presence of highly population-specific protein-truncating variants (PTVs) in human genes using latest population-scale data. We discovered an unexpectedly high incidence of population-specific PTVs in all major ancestral groups. This does not conform to a recently proposed model, indicating either systemic differences in disease penetrance in different human populations, or a failure of current annotation criteria to accurately predict the loss-of-function potential of PTVs. We show that low-confidence pLoF variants are enriched in genes with non-uniform PTV count distribution, and developed a computational tool called LoFfeR that can efficiently predict tolerated pLoF variants. To evaluate the performance of LoFfeR, we use a set of known pathogenic and benign PTVs from the ClinVar database, and show that LoFfeR allows for a more accurate annotation of low-confidence pLoF variants compared to existing methods. Notably, only 4.4% of protein-truncating gnomAD SNPs in canonical transcripts can be filtered out using a recommended threshold value of the recently proposed pext score, while up to 10.9% of such variants are filtered using LoFfeR with the same false positive rate. Hence, we believe that LoFfeR can be used for additional filtering of low-confidence pLoF variants in population genomics and medical genetics studies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 96%
- Mitochondrial DNA variation across 56,434 individuals in gnomAD 95%
- Structural variants are a major source of gene expression differences in humans and often affect multiple nearby genes 95%
Similar papers in this journal
Similar papers in this journal
- TADA - a Machine Learning Tool for Functional Annotation based Prioritisation of Putative Pathogenic CNVs 95%
- Assembly and Annotation of an Ashkenazi Human Reference Genome 94%
- HOPS: a quantitative score reveals pervasive horizontal pleiotropy in human genetic variation is driven by extreme polygenicity of human traits and diseases 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.