TM-Vec 2: Accelerated Protein Homology Detection for Structural Similarity
Keluskar, A.; Batra, P.; Bezshapkin, V.; Morton, J. T.; Zhu, Q.
Show abstract
Understanding protein function is an essential aspect of many biological applications. The exponential growth of protein sequence databases has created a critical bottleneck for structural homology detection. While billions of protein sequences have been identified from sequencing data, the number of protein folds underlying biology is surprisingly limited, likely numbering tens of thousands. The "sequence-fold gap" limits the success of functional annotation methods that rely on sequence homology, especially for newly sequenced, divergent microbial genomes. TM-Vec is a deep learning architecture that can predict TM scores as a metric of structural similarity directly from sequence pairs, bypassing true structural alignment. However, the computational demands of its protein language model (PLM) embeddings create a significant bottleneck for large-scale database searches. In this work, we present two innovations: TM-Vec 2, a new architecture that optimizes the computationally-heavy sequence embedding step, and TM-Vec 2s, a highly efficient model created by distilling the knowledge of the TM-Vec 2 model. Our new models were benchmarked for both accuracy and speed on using the CATH and SCOPe domains for large-scale database queries. We compare them to state-of-the-art models to observe that TM-Vec 2s achieves speedups of up to 258x over the original TM-Vec and 56x over Foldseek for large-scale database queries, while achieving higher accuracy compared to the original TM-Vec model.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Analysis of Protein Language Model Embeddings for Fold Prediction 97%
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 97%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.