Rewriting protein alphabets with language models
Pantolini, L.; Studer, G.; Engist, L.; Pudziuvelyte, I.; Pommerening, F.; Waterhouse, A. M.; Tauriello, G.; Steinegger, M.; Schwede, T.; Durairaj, J.
Show abstract
Detecting remote homology with speed and sensitivity is crucial for tasks like function annotation and structure prediction. We introduce a novel approach using contrastive learning to convert protein language model embeddings into a new 20-letter alphabet, TEA, enabling highly efficient large-scale protein homology searches. Searching with our alphabet performs on par with and complements structure-based methods without requiring any structural information, and with the speed of sequence search. Ultimately, we bring the exciting advances in protein language model representation learning to the plethora of sequence bioinformatics algorithms developed over the past century, offering a powerful new tool for biological discovery.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mapping the space of protein binding sites with sequence-based protein language models 97%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 96%
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 96%
Similar papers in this journal
Similar papers in this journal
- AlphaFold Model Quality Self-Assessment Improvement Via Deep Graph Learning 97%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.