A scalable approach to resolving variants of uncertain significance
Tejura, M.; Chen, Y.; McEwen, A. E.; Stewart, R.; Sverchkov, Y.; Laval, F.; Woo, I.; Zeiberg, D.; Shen, R.; Fayer, S.; Stone, J.; Smith, N.; Casadei, S.; Wang, Z. R.; Snyder, M.; Capodanno, B. J.; Gupta, P.; Benazouz, M.; Jain, S.; Heidl, S.; Muffley, L.; Dong, S.; Lin, K.; Hitz, B. C.; Gabdank, I.; Da, E. Y.; Best, S.; Grindstaff, S.; Reinhart, D.; Rodriguez-Salas, L.; Seid, O.; Vandi, A. J.; Wenman, C.; Wheelock, M. K.; Pendyala, S.; Holmes, D.; Xu, A.; Hosokai, A.; Tixhon, M.; Reno, C.; Ewald, J. D.; Spirohn-Fitzgerald, K.; Teelucksingh, T.; Hao, T.; Chen, Z. S.; Haghighi, M.; Hamid, A. K.;
Show abstract
Over 90% of missense variants across [~]4,000 disease-associated genes are variants of uncertain significance (VUS). Experimental variant effect measurements provide critical evidence about pathogenicity and inform disease biology, but most variants lack data and clinical translation has been limited. The Impact of Genomic Variation on Function Consortium generated experimental data for 62,215 variants across ten genes using multiplexed assays and 1,407 variants across 163 genes using arrayed assays, curated 193,139 additional community-generated variant effect measurements across 30 additional genes, and developed automated calibration methods for translating experimental data and variant effect predictions into clinical evidence. To reduce current VUS, we developed a scalable workflow using only experimental and predictive evidence, enabling reclassification of 75% of the 16,115 VUS in these genes as pathogenic or benign with <1% error. To minimize future VUS, we analyzed >90,000 unobserved variants; 62% had enough evidence to be "preclassified" as pathogenic or benign. We validated our data, evidence and classifications using All of Us and created interactive resources to enable clinical use of the calibrated data. Thus, for 40 genes, representing 1% of the clinical genome, we resolve most existing VUS and future variants, illustrating how systematic use of scalable evidence can empower genomic medicine.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 98%
- Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance 97%
- Tissue-specific enhancer-gene maps from multimodal single-cell data identify causal disease alleles 97%
Similar papers in this journal
Similar papers in this journal
- Large-scale clinical interpretation of genetic variants using evolutionary data and deep learning 98%
- A genome-wide mutational constraint map quantified from variation in 76,156 human genomes 98%
- A familial, telomere-to-telomere reference for human de novo mutation and recombination from a four-generation pedigree 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.