Leveraging protein language models and scoring function for Indel characterisation and transfer learning
Gracia I Carmona, O.; Leipart, V.; Amdam, G. V.; Orengo, C.; Fraternali, F.
Show abstract
1.Protein language models (PLMs) are increasingly used to assess the impact of genetic variation on proteins. By leveraging sequence information alone, PLMs achieve high performance and accuracy and can outperform traditional pathogenicity predictors specifically designed to identify harmful variants contributing to diseases. PLMs can perform zero-shot inference, making predictions without task-specific fine-tuning, offering a simpler and less overfitting-prone alternative to complex methods. However, studying in-frame insertions and deletions (indels) with PLMs remains challenging. Indels alter protein length, making direct comparisons between wildtype and mutant sequences not straightforward. Additionally, indel pathogenicity is less studied than other genetic variants, such as single nucleotide variants, resulting in a lack of annotated datasets. Despite these challenges, approaches that leverage PLMs through transfer learning have emerged, making it possible to capture the features needed for more accurate predictions. Still, the current approaches are limited in terms of allowed organisms, indel length, and interpretability. In this work, we devise a new scoring approach for indel pathogenicity prediction (IndeLLM) that provides a solution for the difference in protein lengths. Our method only uses sequence information and zero-shot inference with a fraction of computing time while achieving performances similar to other indel pathogenicity predictors. We used our approach to construct a simple transfer learning approach for a Siamese network, which outperformed all tested indel pathogenicity prediction methods (Matthews correlation coefficient = 0.77). IndeLLM is universally applicable across species since PLMs are trained on diverse protein sequences. To enhance accessibility, we designed a plug-and-play Google Colab notebook that allows easy use of IndeLLM and visualisation of the impact of indels on protein sequence and structure. The tool is available on GitHub https://github.com/OriolGraCar/IndeLLM and Colab https://colab.research.google.com/drive/1CgwprttaNFR_KeJGyFzP0a0C9Y wc4P. Graphical Abstract, if needed, or logo til include on Google Colab O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=159 SRC="FIGDIR/small/642715v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@e5d1b0org.highwire.dtl.DTLVardef@299b19org.highwire.dtl.DTLVardef@1859f02org.highwire.dtl.DTLVardef@18a7ba9_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 96%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
Similar papers in this journal
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 96%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.