A new paradigm for biological sequence retrieval inspired by natural language processing and database research
Rousseau, A.-J.; Lemal, S.; Korovin, Y.; Triantopoulos, G.; Brands, I.; Biemans, M.; Van Hyfte, D.; Valkenborg, D.
Show abstract
Nearly-exponential growth and heterogeneity of biological sequence data make the task of biological sequence retrieval from databases more important and challenging than ever. In this manuscript, we present a novel search algorithm involving an indexing scheme based on patterns discovered by natural language processing, i.e., short strings of nucleotides or amino acids, akin to standard k-mers, but mined from cumulative cross-species omic data repositories. More specifically, we benchmark the quality of the sequence retrieval process by comparing to BLASTP, a heuristic algorithm for the alignment of genomics or protein sequence data. The main argumentation is that to retrieve biological similar sequences it is not needed to mimic the alignment procedures as it is performed by BLAST. Our results suggests that the HYFT-indexing and searching is a good alternative and a static, alignment-free method to retrieve homologous sequence down to 50% sequence identity.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Smash++: an alignment-free and memory-efficient tool to find genomic rearrangements 96%
- AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data 96%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 96%
Similar papers in this journal
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 95%
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 95%
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.