TM-Vec: template modeling vectors for fast homology detection and alignment
Hamamsy, T.; Morton, J.; Berenberg, D.; Carriero, N.; Gligorijevic, V.; Blackwell, R.; Strauss, C.; Leman, J. K.; Cho, K.; Bonneau, R.
Show abstract
Exploiting sequence-structure-function relationships in molecular biology and computational modeling relies on detecting proteins with high sequence similarities. However, the most commonly used sequence alignment-based methods, such as BLAST, frequently fail on proteins with low sequence similarity to previously annotated proteins. We developed a deep learning method, TM-Vec, that uses sequence alignments to learn structural features that can then be used to search for structure-structure similarities in large sequence databases. We train TM-Vec to accurately predict TM-scores as a metric of structural similarity for pairs of structures directly from sequence pairs without the need for intermediate computation or solution of structures. For remote homologs (sequence similarity [≤] 10%) that are highly structurally similar (TM-score ? 0.6), we predict TM-scores within 0.026 of their value computed by TM-align. TM-Vec outperforms traditional sequence alignment methods and performs similar to structure-based alignment methods. TM-Vec was trained on the CATH and SwissModel structural databases and it has been tested on carefully curated structure-structure alignment databases that were designed specifically to test very remote homology detection methods. It scales sub-linearly for search against large protein databases and is well suited for discovering remotely homologous proteins.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 98%
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 96%
- Evolutionary velocity with protein language models 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.