Back

Unsupervised detection of antimicrobial-resistance determinants by coupling protein-language-models and evolutionary signatures

Aris-Brosou, S.; Kouassi, A. A. M.

2026-08-20 bioinformatics
10.64898/2026.08.17.745273 bioRxiv
Show abstract

Antimicrobial resistance (AMR) is among the most pressing threats to global health, yet our ability to find resistance determinants is largely confined to what reference databases already contain: homology search and supervised classifiers recognize variants of known genes but are, by construction, blind to the larger environmental and clinical reservoir of determinants that have not yet been catalogued. To address this critical limitation, we tested whether resistance determinants can be flagged \emph{without} using any resistance label, by leveraging evolutionary signatures that acquisition and adaptation leave in bacterial genomes. We describe a label-free, multi-view framework that scores every gene family of a pangenome on five orthogonal axes: protein-language-model novelty relative to known protein space, mobility/compositional anomaly, episodic positive selection, presence/ absence homoplasy, and reconciliation-inferred horizontal transfer. These views were then combined based on a conjunctive (weighted geometric-mean) rule, so that only families implicated by several independent lines of evolutionary evidence score high. On a controlled simulation the conjunction recovers all planted determinants where no single view is specific. Applied without retraining to the Escherichia coli (n=150) and Klebsiella pneumoniae (n=150) pangenomes, the results confirm that known determinants are almost entirely accessory and concentrates them near the top of the ranking for K. pneumoniae (6.6-fold enrichment in the top 1%), but not for E. coli. Ablation shows presence/absence homoplasy carries most of the signal, that episodic selection is counterproductive, and that an equal-weighted conjunction is suboptimal. The framework offers a reproducible, database-independent shortlist of candidate determinants and a honest accounting of where evolutionary signal is, but is not sufficient on its own.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.