Unsupervised detection of antimicrobial-resistance determinants by coupling protein-language-models and evolutionary signatures
Aris-Brosou, S.; Kouassi, A. A. M.
Show abstract
Antimicrobial resistance (AMR) is among the most pressing threats to global health, yet our ability to find resistance determinants is largely confined to what reference databases already contain: homology search and supervised classifiers recognize variants of known genes but are, by construction, blind to the larger environmental and clinical reservoir of determinants that have not yet been catalogued. To address this critical limitation, we tested whether resistance determinants can be flagged \emph{without} using any resistance label, by leveraging evolutionary signatures that acquisition and adaptation leave in bacterial genomes. We describe a label-free, multi-view framework that scores every gene family of a pangenome on five orthogonal axes: protein-language-model novelty relative to known protein space, mobility/compositional anomaly, episodic positive selection, presence/ absence homoplasy, and reconciliation-inferred horizontal transfer. These views were then combined based on a conjunctive (weighted geometric-mean) rule, so that only families implicated by several independent lines of evolutionary evidence score high. On a controlled simulation the conjunction recovers all planted determinants where no single view is specific. Applied without retraining to the Escherichia coli (n=150) and Klebsiella pneumoniae (n=150) pangenomes, the results confirm that known determinants are almost entirely accessory and concentrates them near the top of the ranking for K. pneumoniae (6.6-fold enrichment in the top 1%), but not for E. coli. Ablation shows presence/absence homoplasy carries most of the signal, that episodic selection is counterproductive, and that an equal-weighted conjunction is suboptimal. The framework offers a reproducible, database-independent shortlist of candidate determinants and a honest accounting of where evolutionary signal is, but is not sufficient on its own.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrating theory and machine learning to reveal determinants of plasmid copy number 95%
- Deciphering polymorphism in 61,157 Escherichia coli genomes via epistatic sequence landscapes 94%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 94%
Similar papers in this journal
- Interpreting tree ensemble machine learning models with endoR 94%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 94%
- CAPYBARA: A Generalizable Framework for Predicting Serological Measurements Across Human Cohorts 93%
Similar papers in this journal
- Oncodrive3D: Fast and accurate detection of structural clusters of somatic mutations under positive selection 93%
- Signatures of cell death and proliferation in perturbation transcriptomics data - from confounding factor to effective prediction 93%
- Single-Cell Trajectory Inference for Detecting Transient Events in Biological Processes 92%
Similar papers in this journal
- What do we gain when tolerating loss? The information bottleneck wrings out recombination 95%
- GenomegaMap: within-species genome-wide dN/dS estimation from over 10,000 genomes 95%
- Positively twisted: The complex evolutionary history of Reverse Gyrase suggests a non-hyperthermophilic Last Universal Common Ancestor 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.