Protein Structural Alignments From Sequence
Morton, J.; Strauss, C.; Blackwell, R.; Berenberg, D.; Gligorijevic, V.; Bonneau, R.
Show abstract
Computing sequence similarity is a fundamental task in biology, with alignment forming the basis for the annotation of genes and genomes and providing the core data structures for evolutionary analysis. Standard approaches are a mainstay of modern molecular biology and rely on variations of edit distance to obtain explicit alignments between pairs of biological sequences. However, sequence alignment algorithms struggle with remote homology tasks and cannot identify similarities between many pairs of proteins with similar structures and likely homology. Recent work suggests that using machine learning language models can improve remote homology detection. To this end, we introduce DeepBLAST, that obtains explicit alignments from residue embeddings learned from a protein language model integrated into an end-to-end differentiable alignment framework. This approach can be accelerated on the GPU architectures and outperforms conventional sequence alignment techniques in terms of both speed and accuracy when identifying structurally similar proteins.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 97%
- Scoring Protein Sequence Alignments Using Deep Learning 97%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 97%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 96%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 96%
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 95%
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 97%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 97%
- Predicting Affinity Through Homology (PATH): Interpretable Binding Affinity Prediction with Persistent Homology 96%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Prop3D: A Flexible, Python-based Platform for Machine Learning with Protein Structural Properties and Biophysical Data 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.