SSAlign: Ultrafast and Sensitive Protein Structure Search at Scale
Wang, L.; Zhang, X.; Wang, Y.; Xue, Z.
Show abstract
The advent of highly accurate structure prediction techniques such as AlphaFold3 is driving an unprecedented expansion of protein structure databases. This rapid growth creates an urgent demand for novel search tools, as even the current fastest available methods like Foldseek face significant limitations in sensitivity and scalability when confronted with these massive repositories. To meet this challenge, we have developed SSAlign, a protein structure retrieval tool that leverages protein language models to jointly encode sequence and structural information, and adopts a two-stage alignment strategy optimized with multi-GPU and multi-process parallelization. On large-scale datasets such as AFDB50, SSAlign outpaces Foldseek by two to three orders of magnitude in search speed, offering unmatched scalability for high-throughput structural analysis. Compared to Foldseek, SSAlign retrieves substantially more high-quality matches on Swiss-Prot and achieves marked performance improvements on SCOPe40, with relative AUC increases of +20.2% at the family level and +33.3% at the superfamily level, demonstrating significantly enhanced sensitivity and recall. In sum, SSAlign achieves TM-align-comparable accuracy with Foldseek-surpassing speed and coverage, offering an efficient, sensitive, and scalable solution for large-scale structural biology and structure-based drug discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 98%
- A Fast Approach for Structural and Evolutionary Analysis Based on Energetic Profile Protein Comparison 97%
- Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction 96%
Similar papers in this journal
- Rapid and accurate protein structure database search using inverse folding model and contrastive learning 97%
- PymolFold: A PyMOL Plugin for API-driven Structure Prediction and Quality Assessment 96%
- To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map 95%
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
- Scoring Protein Sequence Alignments Using Deep Learning 95%
- OPUS-X: An Open-Source Toolkit for Protein Torsion Angles, Secondary Structure, Solvent Accessibility, Contact Map Predictions, and 3D Folding 95%
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 96%
- Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2 95%
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
- ProteinDJ: a high-performance and modular protein design pipeline 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.