Back

SyMetrics: An Integrated Machine Learning Model for Evaluating the Pathogenicity of Synonymous Variants in the Human Genome

Bundalian, L. T.; Strnadova, M.; Garten, F.; Horn, S.; Stenzel, U.; Lemke, J.; Biskup, S.; Schulte, B.; May, P.; Bosebeck, F.; Garten, A.; Thor, D.; Schulz, A.; Hentschel, J.; Kelso, J.; Schoneberg, T.; Le Duc, D.

2025-03-23 genetic and genomic medicine
10.1101/2025.03.21.25324414 medRxiv
Show abstract

Synonymous single nucleotide variants (sSNVs), traditionally seen as neutral, are now recognized for their biological impact. To assess their relevance, we developed SyMetrics, a framework that integrates predictors of splicing, RNA stability, evolutionary conservation, codon usage, synonymous variation effects, sequence properties, and allele frequency. We analyzed all possible sSNVs across the human genome, and our machine-learning model achieved 97% accuracy in distinguishing deleterious from benign variants, with a ROC-AUC of 0.89, outperforming individual predictors. Our estimates indicate that about 1.98 {+/-} 0.17% of sSNVs absent from population databases are damaging (roughly 900, 000 sSNVs), with an odds ratio of 3.87 for deleteriousness compared to common sSNVs (p < 0.05). To validate predictions, we performed functional assays on selected sSNVs in the AVPR2 gene. In a clinical cohort, we identified 15 predicted deleterious sSNVs in genes linked to patient phenotypes; 9 were classified as (likely) pathogenic while 6 were variants of uncertain significance (VUS) per American College of Medical Genetics guidelines. For three VUS, segregation data supported their suspected inheritance patterns (de novo, X-linked). Our findings underscore the functional importance of sSNVs. To support further research and clinical applications, we provide a Python package and web application for evaluating these variants comprehensively. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=102 SRC="FIGDIR/small/25324414v1_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@12f1627org.highwire.dtl.DTLVardef@57968borg.highwire.dtl.DTLVardef@5cb42borg.highwire.dtl.DTLVardef@38733c_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in NAR Genomics and Bioinformatics (predicted rank #1) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.