Ultra-fast and highly sensitive protein structure alignment with segment-level representations and block-sparse optimization
Litfin, T.; Zhou, Y.; von Itzstein, M.
Show abstract
Deep learning models for protein structure prediction have given rise to extreme growth in 3D structure data. As a result, traditional methods for geometric structure alignment are too slow to effectively search modern structure libraries. In this study we introduce SPfast - a fully geometric method for structure-based alignment which accelerates search by more than 2 orders of magnitude while increasing sensitivity by 21% and 5% compared with foldseek and TMalign respectively. Using the significant speed of SPfast to conduct more than 100B pairwise comparisons between bona fide uncharacterized proteins and a large-scale, annotated structure library uncovers new biological insights relating to type III secretion in pathogenic bacteria and identifies novel toxin-antitoxin systems. Putative SPfast-based functional assignments are supported by orthogonal evidence including shared genomic context and high-confidence AlphaFold3 complex modelling.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 97%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 97%
- AlphaFold2 fails to predict protein fold switching 96%
Similar papers in this journal
- A Suite of Designed Protein Cages Using Machine Learning Algorithms and Protein Fragment-Based Protocols 96%
- GemSpot: A Pipeline for Robust Modeling of Ligands into CryoEM Maps 95%
- DomainFit: Identification of Protein Domains in cryo-EM maps at Intermediate Resolution using AlphaFold2-predicted Models 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.