Amira: gene-space de Bruijn graphs to improve the detection of AMR genes from bacterial long reads
Anderson, D.; Lima, L.; Le, T.; Judd, L. M.; Wick, R. R.; Iqbal, Z.
Show abstract
Accurate detection of antimicrobial resistance (AMR) genes is essential for the surveillance, epidemiology and genotypic prediction of AMR. This is typically done by generating an assembly from the sequencing reads of a bacterial isolate and running AMR gene detection tools on the assembly. However, despite advances in long-read sequencing that have greatly improved the quality and completeness of bacterial genome assemblies, assembly tools remain prone to large-scale errors caused by repeats in the genome, leading to inaccurate detection of AMR gene content and consequent impact on resistance prediction. In this work we present Amira, a tool to detect AMR genes directly from unassembled long-read sequencing data. Amira leverages the fact that multiple consecutive genes lie within a single read to construct gene-space de Bruijn graphs where the k-mer alphabet is the set of genes in the pan-genome of the species under study. Through this approach, the reads corresponding to different copies of AMR genes can be effectively separated based on the genomic context of the AMR genes, and used to infer the nucleotide sequence of each copy. Amira achieves significant improvements in genomic copy number recall and nucleotide accuracy, demonstrated through objective simulations and comparison with alternative read and assembly-based methods on samples with manually curated truth assemblies. Applied to a dataset of 32 Escherichia coli samples with diverse AMR gene content, Amira achieves a mean genomic-copy-number recall of 98.4% with precision 97.9% and nucleotide accuracy 99.9%. Finally, we show that Amira consistently detects more true AMR genes across all E. coli, K. pneumoniae and E. faecium nanopore datasets from the ENA (n=8580, 2448 and 415 respectively) than an assembly-based approach.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DRAM for distilling microbial metabolism to automate the curation of microbiome function 96%
- New candidates for regulated gene integrity revealed through precise mapping of integrative genetic elements 96%
- Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity 95%
Similar papers in this journal
Similar papers in this journal
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 96%
- Taxor: Fast and space-efficient taxonomic classification of long reads with hierarchicalinterleaved XOR filters 96%
- PAN-GWES: Pangenome-spanning epistasis and co-selection analysis via de Bruijn graphs 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.