CARD k-mers: Unmasking the pathogen hosts and genomic contexts of antimicrobial resistance genes in metagenomic sequences
Wlodarski, M. A.; Lau, T. T. Y.; Alcock, B. P.; Raphenya, A. R.; Ta, T. E.; Maguire, F.; Beiko, R. G.; McArthur, A. G.
Show abstract
Antimicrobial resistance (AMR) is a global health crisis requiring rapid surveillance across human, agricultural, and environmental systems. A major challenge during outbreaks is not only detecting antimicrobial resistance genes (ARGs), but also unmasking their pathogen hosts and genomic context, as ARGs alone do not fully capture AMR risk. Pathogen identification is often essential for guiding effective treatment. While culture-based methods remain the diagnostic gold standard, they are slow and sometimes impractical. Faster metagenomic (mNGS) tools typically detect either ARGs, taxonomy, or genomic context, but rarely all three, resulting in fragmented surveillance. Existing k-mer classifiers like Kraken2 and CLARK, designed for general taxonomy, often perform poorly on AMR-specific sequences. We introduce CARD k-mers, the first tool built to jointly predict species-level taxonomy and genomic context (plasmid vs. chromosome) for ARGs in short metagenomic reads. Integrated with the Comprehensive Antibiotic Resistance Database (CARD), CARD k-mers enables rapid, context-aware assignment of ARGs to their likely pathogen and genomic element origin. In benchmarking with 103,456 in-silico pathogen-specific AMR alleles, CARD k-mers outperformed Kraken2 and CLARK by 10.85% and 15.2%, respectively, and correctly classified the genomic context of 4,590 chromosome- and 176 plasmid-specific ARGs. The tool operates at speeds exceeding 675,000 metagenomic reads per minute. By delivering fast, accurate, and context-rich classification of ARGs, CARD k-mers significantly advances untargeted AMR surveillance and is accessible to users with basic command-line experience for use in both clinical and environmental pipelines. CARD k-mers is available at: https://github.com/arpcard/rgi.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 96%
- Comparison of R9.4.1/Kit10 and R10/Kit12 Oxford Nanopore flowcells and chemistries in bacterial genome reconstruction 96%
- K-mer based prediction of Clostridioides difficile relatedness and ribotypes 95%
Similar papers in this journal
- A comparison of short- and long-read whole genome sequencing for microbial pathogen epidemiology 95%
- A method to correct for local alterations in DNA copy number that bias functional genomics assays applied to antibiotic-treated bacteria 94%
- Nanopore adaptive sampling effectively enriches bacterial plasmids 94%
Similar papers in this journal
- Unveiling the Microbial Realm with VEBA 2.0: A modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic, and viral multi-omics from either short- or long-read sequencing 95%
- Automating microbial taxonomy workflows with PHANTASM: PHylogenomic ANalyses for the TAxonomy and Systematics of Microbes 95%
- Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity 95%
Similar papers in this journal
- Real-time Plasmid Transmission Detection Pipeline 96%
- Development of an amplicon nanopore sequencing strategy for detection of mutations conferring intermediate resistance to vancomycin in Staphylococcus aureus strains 96%
- GenomicGapID: Leveraging Spatial Distribution of Conserved Genomic Sites for Broad-Spectrum Microbial Identification 95%
Similar papers in this journal
- SYNTERUPTOR: mining genomic islands for non-classical specialised metabolite gene clusters 94%
- A primer-independent DNA polymerase-based method for competent whole-genome amplification of intermediate to high GC sequences 94%
- Life at the extremes: Maximally divergent microbes with similar genomic signatures linked to extreme environments 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.