ProtSEC: Ultrafast Protein Sequence Embedding in Complex Space Using Fast Fourier Transform
Raju, R. S.; Islam, R.
Show abstract
Among the various approaches for representing protein sequences as vectors, embeddings derived from protein language models (PLMs) have been empirically shown to enhance accuracy in downstream bioinformatics tasks. However, the substantial computational demands of PLMs, both during training and inference, pose significant challenges. We introduce ProtSEC (Protein Sequence Embedding in Complex Space), a novel approach that encodes each amino acid as a unique complex number derived from the BLOSUM62 substitution matrix. By modeling protein sequences as complex signals and applying the Fast Fourier Transform (FFT), ProtSEC generates embeddings in the complex space. Unlike PLMs, ProtSEC requires no pre-training on large protein sequence datasets and operates independently of any pre-trained models. Our benchmarking demonstrate that ProtSEC achieves a 20,000-fold reduction in runtime and an 85-fold improvement in memory efficiency compared to popular PLMs (e.g., esm2_3B, esm2_35M, prot_t5, prot_bert). Depending on the task, ProtSEC demonstrates either superior or comparable accuracy to PLMs in sequence similarity search, sequence classification and phylogenetic tree reconstruction. ProtSEC provides fast and accurate protein sequence embeddings in complex numbers, facilitating efficient integration into diverse downstream bioinformatics workflows. ProtSEC is available at https://github.com/omics-lab/ProtSEC.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 97%
- SPAED: Harnessing AlphaFold Output for Accurate Segmentation of Phage Endolysin Domains 95%
- CICLOP: A Robust, Faster, and Accurate Computational Framework for Protein Inner Cavity Detection 95%
Similar papers in this journal
- Evaluating the Significance of Embedding-Based Protein Sequence Alignment with Clustering and Double Dynamic Programming for Remote Homology 97%
- ProteinGLUE: A multi-task benchmark suite for self-supervised protein modeling. 97%
- Two sequence- and two structure-based ML models have learned different aspects of protein biochemistry 96%
Similar papers in this journal
- To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map 96%
- Rapid and accurate protein structure database search using inverse folding model and contrastive learning 95%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 95%
Similar papers in this journal
Similar papers in this journal
- TemBERTure: Advancing protein thermostability prediction with Deep Learning and attention mechanisms 95%
- Nucleotide augmentation for machine learning-guided protein engineering 94%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.