Predicting Genetic Variant Pathogenicity Using Vector Embeddings
Wu, J.; Muriello, M.; Basel, D. G.; Gai, X.
Show abstract
BackgroundInterpreting the pathogenicity of genetic variants remains a critical bottleneck in genomic medicine. Millions of variants of uncertain significance (VUS) hinder the clinical application of genetic findings. Traditional computational approaches often rely on hand-engineered features and fail to fully capture the complexity of multidimensional genomic annotations. MethodsWe developed a novel semantic embedding framework, VUS.Life, that transforms variant annotations into natural language descriptions and leverages pre-trained language models to capture pathogenicity-relevant relationships in high-dimensional vector space. This approach enables direct pathogenicity prediction through representation learning rather than traditional feature-based methods. ResultsWe evaluated the framework using curated variants from three disease-associated genes: BRCA1 (n=3,311) and BRCA2 (n=4,074) from BRCA Exchange, and FBN1 (n=1,532) from ClinVar. Variant annotations from the Variant Effect Predictor (VEP) were converted into natural language while excluding fields linked to known pathogenicity assertions to prevent data leakage. We then embedded these descriptions using three models: MPNet (all-mpnet-base-v2), Googles text-embedding-004, and MedEmbed-large-v0.1. A k-nearest neighbor (k-NN) approach (up to 20 neighbors) was used to predict pathogenicity for new or unreviewed variants. Dimensionality reduction techniques (PCA, t-SNE, UMAP) enabled visualization of the embedding spaces. k-NN classification showed exceptional performance across all genes and embedding models. For BRCA1, overall accuracy ranged from 97.3% (Google) to 97.9% (MPNet), with benign/likely benign accuracy of 95.1-97.2% and pathogenic/likely pathogenic accuracy of 97.9-98.4%. For BRCA2, overall accuracy ranged from 97.9% (Google) to 99.1% (MedEmbed), with benign/likely benign accuracy of 96.6-99.2% and pathogenic/likely pathogenic accuracy of 98.6-99.4%. FBN1 validation confirmed the methods generalizability, with accuracy exceeding 96% across all embeddings. Application to not-yet-reviewed BRCA1/2 variants demonstrated the frameworks practicality and scalability, with unknown variants aligning closely to known benign or pathogenic clusters. ConclusionsThis semantic embedding framework, VUS.Life, accurately captures pathogenicity-relevant features from complex variant annotations, enabling high-accuracy (>96%) automated classification across multiple genes and models. The approach generalizes beyond well-curated genes and supports scalable, interpretable, and representation-based classification of VUS. It holds significant promise for alleviating the variant interpretation bottleneck in clinical genomics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AVADA Enables Automated Genetic Variant Curation Directly from the Full Text Literature 95%
- ClinPhen extracts and prioritizes patientphenotypes directly from medical records to accelerate genetic disease diagnosis 95%
- InpherNet provides attractive monogenic disease gene hypotheses using patient genes indirect neighbors 95%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 95%
- Improving Automated Deep Phenotyping Through Large Language Models Using Retrieval Augmented Generation 94%
Similar papers in this journal
- Evidence-based calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for clinical use of PP3/BP4 criteria 95%
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 94%
- Interpretable Clinical Genomics with a Likelihood Ratio Paradigm 94%
Similar papers in this journal
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 94%
- RNAIndel: a machine-learning framework for discovery of somatic coding indels using tumor RNA-Seq data 94%
- Bayesian Estimation of Allele-Specific Expression in the Presence of Phasing Uncertainty 94%
Similar papers in this journal
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 93%
- Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders 93%
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.