Back

Sequence alignment using large protein structure alphabets doubles sensitivity to remote homologs

Edgar, R. C.

2024-05-27 bioinformatics
10.1101/2024.05.24.595840 bioRxiv
Show abstract

Recent breakthroughs in protein fold prediction from amino acid sequences have unleashed a deluge of new structures, raising new opportunities for expanding insights into the universe of proteins and pursuing practical applications in bio-engineering and therapeutics while also presenting new challenges to protein search and analysis algorithms. Here, I describe Reseek, a protein alignment algorithm which improves sensitivity in protein homolog detection compared to state-of-the-art methods including DALI, TM-align and Foldseek, with improved speed over Foldseek, the fastest previous method. Reseek is based on alignment of sequences where each residue in the protein backbone is represented by a letter in a novel "mega-alphabet" of 85,899,345,920 ([~] 1011) distinct states. Code is available at https://github.com/rcedgar/reseek.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.