Back

ProtSEC: Ultrafast Protein Sequence Embedding in Complex Space Using Fast Fourier Transform

Raju, R. S.; Islam, R.

2025-08-22 bioinformatics
10.1101/2025.08.17.670693 bioRxiv
Show abstract

Among the various approaches for representing protein sequences as vectors, embeddings derived from protein language models (PLMs) have been empirically shown to enhance accuracy in downstream bioinformatics tasks. However, the substantial computational demands of PLMs, both during training and inference, pose significant challenges. We introduce ProtSEC (Protein Sequence Embedding in Complex Space), a novel approach that encodes each amino acid as a unique complex number derived from the BLOSUM62 substitution matrix. By modeling protein sequences as complex signals and applying the Fast Fourier Transform (FFT), ProtSEC generates embeddings in the complex space. Unlike PLMs, ProtSEC requires no pre-training on large protein sequence datasets and operates independently of any pre-trained models. Our benchmarking demonstrate that ProtSEC achieves a 20,000-fold reduction in runtime and an 85-fold improvement in memory efficiency compared to popular PLMs (e.g., esm2_3B, esm2_35M, prot_t5, prot_bert). Depending on the task, ProtSEC demonstrates either superior or comparable accuracy to PLMs in sequence similarity search, sequence classification and phylogenetic tree reconstruction. ProtSEC provides fast and accurate protein sequence embeddings in complex numbers, facilitating efficient integration into diverse downstream bioinformatics workflows. ProtSEC is available at https://github.com/omics-lab/ProtSEC.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.