Snekmer Learn/Apply: A kmer-based vector similarity approach to proteinclassification suitable for metagenomic datasets
Nitka, T. A.; Jacobson, J.; Chang, C. H.; Krause, G. R.; Wheeler, T. J.; Egbert, R. G.; Nelson, W. C.; McDermott, J. E.
Show abstract
Advances in whole genome sequencing have led to a rapid and ongoing increase in the amount of sequence data available, but 40-50% of known genes have no functional annotation and only 25-30% have specific functional annotations. Current functional annotation approaches typically rely on computationally expensive pairwise or multiple sequence alignments, preventing rapid development of models for novel protein functions and sometimes limiting methods to one ontology. Representation of sequence in short segments (kmers) has been used in many applications for nucleotide sequence, and more recently has been applied to protein sequence as well. We previously developed Snekmer, a tool which uses kmer patterns to develop alignment-free individual protein family models. Other approaches, such as MMSeqs2 and DIAMOND, use protein kmers as a fast filter to reduce search space for subsequent sequence alignment. Here, we describe a novel addition to the Snekmer tool which builds kmer libraries for protein families and uses those libraries to map functional annotations to new sequences. We first demonstrate that our method accurately applies TIGRFAMs annotations to protein fragments and to a low-sequence identity benchmark dataset, and further use it to annotate a set of drought stress associated soil and rhizosphere metagenome sequences with higher sensitivity towards several important protein function classes than that shown by HMMs. We have incorporated this workflow into Snekmer.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- An nf-core framework for the systematic comparison of alternative modeling tools: the multiple sequence alignment case study 95%
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 94%
- ganon2: up-to-date and scalable metagenomics analysis 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.