Back

Bioactivity-Driven Prediction of Antibacterial Synergy Using Machine Learning Models

Yousefabadi, H.; Mehrmohamadi, M.

2025-12-16 bioinformatics
10.64898/2025.12.14.694088 bioRxiv
Show abstract

MotivationPredicting antibacterial drug synergy remains difficult due to strain variability and the limited scale of experimentally tested combinations. Existing machine-learning approaches often rely on permissive cross-validation schemes that allow drug pairs to appear across folds, inflating performance. A rigorous evaluation framework and scalable feature representation are needed for robust generalization. ResultsWe assembled a curated dataset of 3,160 drug-pair-strain interactions covering 97 compounds and 10 bacterial strains. We then developed HALO (Held-out Antibiotic interaction Learning from latent bioactivity Observations), a synergy-prediction framework in which each drug pair is encoded using multi-level Chemical Checker (CC) similarity features spanning chemical, target, network, cellular, and clinical bioactivity domains. Under strictly nested, pair-heldout cross-validation (CV1), HALO achieved stable generalization to unseen combinations (accuracy {approx} 0.75; ROC-AUC = 0.82). Performance depended strongly on evaluation stringency: models performed well under random splits but degraded when required to generalize to unseen drug pairs and strain contexts. Despite these constraints, HALO generalized to an independent set of Loewe- measurements, achieving ROC-AUC = 0.85 for distinguishing synergy from antagonism. These results demonstrate that multi-level bioactivity signatures provide a scalable, interpretable basis for predicting antibacterial synergy and reveal the performance limits of current models under rigorous evaluation. Availability and ImplementationCode, data-processing scripts, and trained models will be available at GitHub repo. Contactmehrmohamadi@ut.ac.ir Supplementary informationSupplementary figures and additional evaluation details are available online.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.