Evaluating Protein Language Model Embeddingsfor Structural Similarity in the Protein-Sequence Twilight Zone
DANWADA, S.; UDOMPRASERT, P.; RAYCHAWDHARY, N.; SEALS, C. D.; Wu, L.; Bhattacharya, S.
Show abstract
Evaluating protein sequence similarity remains challenging in the protein-sequence twilight zone (20-35% sequence identity), where traditional methods often fail. In this study, we evaluate whether mean-pooled embeddings from four protein language models: ESM-1b, ESM-2, ProtT5, and ProstT5 can estimate pairwise structural similarity without performing sequence alignment. The benchmark dataset includes 20,445 PISCES protein pairs with sequence identity [≤]30%, representing the protein-sequence twilight zone, with TM-align-derived TMmin used as the structural ground truth. Protein embeddings are compared using cosine similarity, Euclidean- and Manhattan-derived similarities, an RBF kernel, and dot product. Among these similarity metrics, cosine similarity performs best across all four models. Moreover, ProstT5 achieves the highest Spearman correlation with TMmin, followed by ESM-2, ProtT5, and ESM-1b, while all four PLMs outperform BLASTP overall. Furthermore, the advantage of PLM embeddings is most pronounced for protein pairs with the lowest sequence identity. ProstT5 also provides the best discrimination between structurally similar and dissimilar protein pairs. Moreover, it offers a favorable balance between similarity performance and the computational requirements of residue-level embedding generation and storage. Overall, these findings support PLM embeddings as an effective alignment-free approach for detecting structural relationships among proteins in the twilight zone.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 95%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 95%
- Improved model quality assessment using sequence and structural information by enhanced deep neural networks 95%
Similar papers in this journal
- From Atoms to Fragments: A Coarse Representation for Functional and Efficient Protein Design 95%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 95%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 95%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.