Back

amR: an R package suite to predict antimicrobial resistance in bacterial pathogens

Ghosh, A.; Brenner, E. P.; Boyer, E. A.; McKim, A. P.; Vang, C. K.; Wolfe, E. P.; Mayer, D. A.; Lesiyon, R. L.; Ravi, J.

2026-07-13 bioinformatics
10.64898/2026.07.10.734579 bioRxiv
Show abstract

MotivationIdentifying bacterial antimicrobial resistance (AMR) is critical for diagnostics and treatment, but resistance is a complex trait arising from myriad mechanisms spanning multiple molecular scales. Existing computational approaches often function as black boxes and rarely explore cross-species or multi-drug patterns. We developed amR, an integrated R package suite that provides a complete framework from bacterial genome data curation to interpretable AMR predictions, enabling identification of resistance mechanisms across species and drugs. ResultsThe amR R package suite contains three modular packages. amRdata downloads genomes and paired antimicrobial susceptibility testing data from BV-BRC and processes them, constructs pangenomes, and extracts features at gene/protein cluster, protein domain, annotated Clusters of Orthologous Groups and ResFinder AMR-associated features, and structural variant scales; data are stored in memory-efficient formats (Parquet, DuckDB). amRml trains interpretable machine learning models per species-drug combination, calculates feature importance and performance metrics, and provides rich ground for hypothesis generation and mechanism discovery. amRviz provides an interactive Shiny dashboard to explore metadata distributions and model performance across species and drugs, visualize top predictive AMR features, and analyze cross-model patterns across geographic/temporal strata. We apply the suite to Shigella sonnei, achieving a median Matthews Correlation Coefficient of 0.89 across 23 drugs and drug classes. With thousands of genomes, multi-scale features, and interpretable models, amR provides an accessible, comprehensive framework for AMR research. The amR package suite is installable via GitHub (https://github.com/JRaviLab/amR; BSD-3-Clause license).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
14.8%
2
PLOS Computational Biology
1863 papers in training set
Top 3%
12.4%
3
Bioinformatics Advances
203 papers in training set
Top 0.2%
11.7%
4
BMC Bioinformatics
457 papers in training set
Top 1%
7.1%
5
Microbial Genomics
225 papers in training set
Top 0.7%
5.1%
50% of probability mass above
6
Nature Communications
5641 papers in training set
Top 29%
5.1%
7
Genome Biology
637 papers in training set
Top 3%
4.0%
8
Nucleic Acids Research
1281 papers in training set
Top 6%
3.2%
9
PeerJ
308 papers in training set
Top 3%
3.1%
10
GigaScience
212 papers in training set
Top 1%
2.7%
11
npj Antimicrobials and Resistance
11 papers in training set
Top 0.1%
2.6%
12
Molecular Biology and Evolution
542 papers in training set
Top 3%
2.1%
13
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.0%
14
PLOS ONE
5266 papers in training set
Top 50%
1.7%
15
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.4%
16
Genome Medicine
183 papers in training set
Top 3%
1.3%
17
Cell Systems
201 papers in training set
Top 3%
1.3%
18
Database
61 papers in training set
Top 0.7%
1.1%
19
eLife
5828 papers in training set
Top 58%
1.1%
20
F1000Research
88 papers in training set
Top 4%
0.8%
21
Scientific Reports
3612 papers in training set
Top 75%
0.8%
22
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
0.8%
23
BMC Genomics
406 papers in training set
Top 8%
0.8%
24
mSystems
394 papers in training set
Top 6%
0.8%
25
Nature Methods
385 papers in training set
Top 7%
0.6%