Back

Fast protein database as a service with kAAmer

Deraspe, M.; Boisvert, S.; Laviolette, F.; Roy, P. H.; Corbeil, J.

2020-04-02 bioinformatics
10.1101/2020.04.01.019984 bioRxiv
Show abstract

Identification of proteins is one of the most computationally intensive steps in genomics studies. It usually relies on aligners that dont accommodate rich information on proteins and require additional pipelining steps for protein identification. We introduce kAAmer, a protein database engine based on amino-acid k-mers, that supports fast identification of proteins with complementary annotations. Moreover, the databases can be hosted and queried remotely.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.