Back

A Systematic Benchmark of Antibiotic Resistance Gene Detection Tools for Shotgun Metagenomic Datasets

Tiwari, S. K.; Ponsero, A. J.; Talas, J.; Grimes, K. P.; Haynes, S.; Telatin, A.

2026-02-06 bioinformatics
10.64898/2026.02.04.703716 bioRxiv
Show abstract

1.Accurate detection of antimicrobial resistance genes (ARGs) from metagenomic data is essential for understanding resistance dissemination within microbial communities, yet tool performance remains influenced by sequencing coverage, community complexity, and dataset variability. In this study, we systematically benchmarked five widely used read-based ARG detection tools (ARGprofiler, KARGA, ARIBA, GROOT, and SRST2) across simulated metagenomic datasets representing varying sequencing coverages, microbial complexities, and approximate realistic metagenomic dataset. The results demonstrated that sequencing coverage is a major determinant of ARG detection accuracy, with reliable detection achieved at 10x coverage and performance stabilizing between 20x and 30x. ARGprofiler exhibited the highest overall F1-score (0.891 at [≥]10x), whereas KARGA showed higher recall at low coverage levels, but lower precision compared to ARGprofiler. Increasing community complexity led to a decline in accuracy across all tools, and under realistic uneven coverage, performance variability increased substantially, with KARGA achieving the highest mean F1- score (0.122 {+/-} 0.067). Runtime evaluation further revealed substantial differences in computational efficiency, with ARGprofiler, SRST2, and GROOT being the most resource-efficient, while KARGA imposed the highest computational burden. Collectively, these findings highlight that both sequencing coverage and community complexity profoundly shape ARG detection outcomes and that tool selection should balance accuracy with computational efficiency. The study also emphasizes the need for standardized benchmarking datasets that reflect true metagenomic complexity to ensure robust and comparable ARG surveillance across analytical pipelines.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.