Fast protein database as a service with kAAmer
Deraspe, M.; Boisvert, S.; Laviolette, F.; Roy, P. H.; Corbeil, J.
Show abstract
Identification of proteins is one of the most computationally intensive steps in genomics studies. It usually relies on aligners that dont accommodate rich information on proteins and require additional pipelining steps for protein identification. We introduce kAAmer, a protein database engine based on amino-acid k-mers, that supports fast identification of proteins with complementary annotations. Moreover, the databases can be hosted and queried remotely.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 93%
- BD5: an open HDF5-based data format to represent quantitative biological dynamics data 93%
- Advancing clinical cohort selection with genomics analysis on a distributed platform 93%
Similar papers in this journal
- Sensitive and error-tolerant annotation of protein-coding DNA with BATH 95%
- LMPred: Predicting Antimicrobial Peptides Using Pre-Trained Language Models and Deep Learning 93%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.