K-MARVEL: K-Mer based Antimicrobial Resistance Virtual Exploration Lab
Mahar, N. S.; Branders, S.; Grabherr, M.; Gupta, I.; Ahmad, R.
Show abstract
The rapid global spread of antimicrobial resistance (AMR) necessitates a new generation of computational tools for its surveillance. While next-generation sequencing offers unprecedented insight into the resistome, current methods face a trade-off: assembly-based approaches are computationally expensive and struggle with complex metagenomes, whereas direct-mapping of long reads is hampered by high error rates that obscure critical resistance-conferring mutations. Here, we present K-MARVEL (K-Mer based Antimicrobial Resistance Virtual Exploration Lab), a novel, open-source method to capture ARGs and resistance-conferring mutations from short and long-read sequencing datasets. It operates in protein k-mer space, providing inherent tolerance to nucleotide-level sequencing errors. On a comprehensive benchmark of 61 long and 49 short-read diverse datasets, K-MARVEL demonstrated superior accuracy, achieving F1-scores of 0.9783 and 0.9754 for short and long-read datasets, respectively. Its implementation in Rust enables high speed through parallelization while guaranteeing memory safety. Computationally, it demonstrated superior performance to conventional assembly-based methods, achieving an average speed up of 7x on short-read datasets and 5x on long-read datasets. In terms of memory footprint, it outperformed the assembly-based approaches for short-read datasets, but its memory footprint was comparable for long-read datasets. Notably, K-MARVEL accurately reconstructs functional genes from genomically fragmented evidence, providing a more comprehensive resistome assessment. In conclusion, K-MARVEL provides a scalable, flexible and memory-efficient solution for AMR surveillance. Its unique capabilities for handling noisy long-read data and complex genomic scenarios make it a powerful tool for researchers and public health scientists. K-MARVEL is open-source and freely available at https://bitbucket.org/amr-avenger/k-marvel under the GPL version 3 license.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing 95%
- cognac: rapid generation of concatenated gene alignments for phylogenetic inferencefrom large whole genome sequencing datasets 95%
- MTG-Link: leveraging barcode information from linked-reads to assemble specific loci 95%
Similar papers in this journal
- Pangenome databases provide superior host removal and mycobacteria classification from clinical metagenomic data 95%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 95%
Similar papers in this journal
- Metagenomics-Toolkit: The Flexible and Efficient Cloud-Based Metagenomics Workflow featuring Machine Learning-Enabled Resource Allocation 95%
- ganon2: up-to-date and scalable metagenomics analysis 95%
- ResistoXplorer: a web-based tool for visual, statistical and exploratory data analysis of resistome data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.