Back

A comprehensive AMR genotype-phenotype database (CABBAGE)

Dickens, E.; Derelle, R.; Beardmore, R. E.; Suresh, A.; Uplekar, S.; Azov, A.; Gurbich, T. A.; El Houdaigui, B.; Keatley, J.; Ochkalova, S.; Koci, O.; Rahman, N. M.; Shivalikanjli, A.; Winterbottom, A.; Yordanova, G.; Parkinson, H.; Yates, A. D.; Finn, R. D.; Lees, J.; Chindelevitch, L.

2026-01-28 genomics
10.1101/2025.11.12.688105 bioRxiv
Show abstract

Addressing the growing threat of antimicrobial resistance (AMR) requires the development of large-scale resources that link bacterial genomic data with phenotypic antimicrobial resistance profiles. Such datasets are essential for advancing genotype-based predictions of resistance to uncover novel resistance mechanisms, as well as identifying and tracking global trends. Here, we describe the development of the Comprehensive Assessment of Bacterial-Based AMR prediction from GEnotypes (CABBAGE) database, linking bacterial genomes to associated antibiotic susceptibility data and relevant metadata across WHO Bacterial Priority Pathogens, sourced from both publications and existing databases, and curated into a format that is compatible with, and extends, both NCBI and ENA formats. The resulting CABBAGE database, comprising over 170,000 unique sequenced isolates and approximately 1.7 million genome-phenotype pairs linked to extensive metadata, represents the largest database of its kind, consolidating existing AMR phenotype-genotype data into a single unified format. CABBAGE encompasses a broad range of antimicrobials, facilitating the analysis of global resistance trends as well as benchmarks of genotype-to-phenotype predictive methods, and empowering further research uses. The database is freely accessible at https://www.ebi.ac.uk/amr and is currently being integrated with the BioSample database, enabling easy access for the AMR research community.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.