Enhancing Clinical Classification of Protein Variants using ESM2 and UMAP
Ugo, L.; Veltri, P.; Guzzi, P. H.
Show abstract
Protein sequences may vary due to mutations in their coding DNA sequence, leading to differences in structure and function. The same protein may exist in multiple variant forms, each potentially leading to distinct phenotypic consequences depending on how the alterations affect its structure, function, or expression. Missense variants are single nucleotide substitutions in the DNA sequence that result in the replacement of one amino acid with another in the corresponding protein, potentially altering its structure, stability, or function. The clinical interpretation of missense variants in protein-coding regions remains a fundamental challenge in genomic medicine. Recent advances in protein language models and manifold learning provide new opportunities for unsupervised extraction of biologically relevant information from protein sequences. In this work, we integrate representations derived from ESM2 (spiegare) with nonlinear dimensionality reduction via UMAP (spiegare) to improve the classification of variants of uncertain significance (VUS) in disease-associated proteins. Our results suggest that this approach improves separability of benign and pathogenic variants, offering a scalable and interpretable strategy for variant prioritization in precision medicine.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exome 95%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 94%
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 93%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 93%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 93%
- VAREANT : a bioinformatics application for gene variant reduction and annotation 93%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 93%
- Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework 93%
- COMIC: Explainable Drug Repurposing via Contrastive Masking for Interpretable Connections 93%
Similar papers in this journal
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 94%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 93%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 93%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 94%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 93%
- Mutation severity spectrum of rare alleles in the human genome is predictive of disease type 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.