Back

Genomic Context as a Predictor of Multidrug Resistance in African Klebsiella pneumoniae: A Feasibility Study with Leave-One-Country-Out Validation

Ahmad, A.; Busair, E.-k.

2026-08-26 bioinformatics
10.64898/2026.08.25.747089 bioRxiv
Show abstract

Multidrug-resistant (MDR) Klebsiella pneumoniae is a leading cause of healthcare-associated mortality in Africa, yet genomic prediction of resistance has relied almost exclusively on resistance-gene detection validated under random data splits. Whether genomic context lineage, capsule and O-locus background, and virulence loci, with all resistance determinants excluded can predict aggregate MDR status, and whether such signal survives geographic transport, remains untested. As a feasibility study, we built an explainable machine-learning framework with leave-one-country-out (LOCO) cross-validation. Phenotypic linkage proved extremely scarce: only 231 of 9,505 strict-K. pneumoniae African NCBI records (2.43%) carry submitter-supplied antibiograms, necessitating a rule-based genotypic MDR proxy label. Country-sufficiency analysis showed LOCO is feasible on the current snapshot 8 countries at n >= 200 genomes but not on any previously published cohort. In a stratified pilot (175 species-confirmed genomes, 9 countries), tree ensembles reached pooled AUROC 0.85 under random splitting but only 0.66 - 0.69 under LOCO; this ~0.15 AUROC geographic-generalization gap suggests that pooled accuracy overstates transportability, though at pilot fold sizes (n <= 20 test genomes) confidence intervals are wide and overlapping. SHAP attributions implicated the ybt virulence locus and O-serotype background, indicating models exploit lineage-associated population structure. The substantive contribution is a leakage-controlled, fully reproducible pipeline indicating that geographic validation, not pooled accuracy, is the operative test for genomic AMR surveillance models.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.