FIND: a software tool for identifying population-enriched pathogenic variants in gnomAD
Horowitz, A. L.; Liebman, A. Z.; Liebman, S. W.
Show abstract
Founder mutations are variants that arose in a single ancestor and became enriched in a descendant population through a bottleneck and endogamy. Identification of pathogenic founder mutations has facilitated efficient targeted screening. More broadly, even without confirmed founder status, identifying pathogenic variants that are enriched within specific populations reveals population-specific disease burden. However, many such variants remain hidden in plain sight within existing datasets. To address this gap, we developed FIND (Founder candidates hidden IN Data), a web tool that identifies pathogenic, likely pathogenic, and predicted loss-of-function variants in gnomAD with frequencies >0.00008 in one ancestry group and at least tenfold higher than in all others (after zeroing populations with four or fewer observed alleles). Testing FIND on the genes FLNC, TMEM127, MYH7, and BRCA2 confirmed its utility and functionality by identifying nine well-known founder mutations and seven candidate founders. Candidates enriched in African American and admixed American populations were validated with the All of Us database, highlighting the utility of this approach for populations historically underrepresented in genetic studies. Source code is freely available at https://github.com/aacoder105/FIND under an MIT license, with a web interface at https://ethnic-variant-mutation-finder.onrender.com/.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Curated Multiple Sequence Alignment for the Adenomatous Polyposis Coli (APC) Gene and Accuracy of In Silico Pathogenicity Predictions 94%
- Assessing the performance of genome-wide association studies for predicting disease risk 93%
- Novel candidates of pathogenic variants of the BRCA1 and BRCA2 genes in a 3,552 Japanese whole-genome sequence dataset (3.5KJPNv2) 93%
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 93%
- Genome-wide survey of tandem repeats by nanopore sequencing shows that disease-associated repeats are more polymorphic in the general population 92%
- Identification of allele-specific KIV-2 repeats and impact on Lp(a) measurements for cardiovascular disease risk 91%
Similar papers in this journal
- Beyond gene-disease validity: capturing structured data on inheritance, allelic-requirement, disease-relevant variant classes, and disease mechanism for inherited cardiac conditions 92%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 92%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 92%
Similar papers in this journal
- Limitations in next-generation sequencing-based genotyping of breast cancer polygenic risk score loci 93%
- High-resolution population-specific recombination rates and their effect on phasing and genotype imputation 92%
- Variable Number Tandem Repeats (VNTRs) as modifiers of breast cancer risk in carriers of BRCA1 185delAG 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.