Back

A new paradigm for biological sequence retrieval inspired by natural language processing and database research

Rousseau, A.-J.; Lemal, S.; Korovin, Y.; Triantopoulos, G.; Brands, I.; Biemans, M.; Van Hyfte, D.; Valkenborg, D.

2023-11-09 bioinformatics
10.1101/2023.11.07.565984 bioRxiv
Show abstract

Nearly-exponential growth and heterogeneity of biological sequence data make the task of biological sequence retrieval from databases more important and challenging than ever. In this manuscript, we present a novel search algorithm involving an indexing scheme based on patterns discovered by natural language processing, i.e., short strings of nucleotides or amino acids, akin to standard k-mers, but mined from cumulative cross-species omic data repositories. More specifically, we benchmark the quality of the sequence retrieval process by comparing to BLASTP, a heuristic algorithm for the alignment of genomics or protein sequence data. The main argumentation is that to retrieve biological similar sequences it is not needed to mimic the alignment procedures as it is performed by BLAST. Our results suggests that the HYFT-indexing and searching is a good alternative and a static, alignment-free method to retrieve homologous sequence down to 50% sequence identity.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.