Back

Species- and Topic-aware Representation Learning for Antimicrobial Peptide Discovery

Padi, S.; Mondal, K.; Kaur, N.; Hoogerheide, D. P.; Heinrich, F.; Mihailescu, E.; Klauda, J. B.; Cardone, A.; Keyrouz, W.

2026-06-01 bioinformatics
10.64898/2026.05.28.728246 bioRxiv
Show abstract

Antimicrobial resistance poses a major global health challenge, necessitating efficient strategies to discover potent antimicrobial peptides (AMPs). While recent generative models can produce many candidate sequences, experimentally validating all generated peptides in wet labs is impractical due to the high costs and time involved in such measurements. As a result, there is a strong demand for accurate predictions of peptide efficacy, typically measured as the minimum inhibitory concentration (MIC). We introduce STAMP, a framework for Species- and Topic-aware Representation Learning in AMP Discovery. This unified machine learning framework allows for cross-species predictions of AMP activity. STAMP integrates protein language model embeddings with species conditioning and topic-aware representations that capture sequence-level patterns, enabling generalizable predictions across multiple bacterial species within a single model. We evaluated STAMP on three benchmark datasets, which include two previously published datasets and a newly curated dataset derived from DBAASP, addressing duplicates and inconsistencies systematically. STAMP achieved strong predictive performance across these datasets, demonstrating a Pearson correlation coefficient (PCC) of 0.837 and an R2 of 0.70, outperforming several baseline models. Importantly, we further validated our prediction model using peptides that were experimentally tested for their antimicrobial activity against E.coli. and S.epidermidis bacteria, demonstrating its real-world applicability. Furthermore, residue-level importance analyses provide insights into the sequence determinants governing antimicrobial activity. Together, these results establish STAMP as a scalable framework for MIC prediction and an effective computational tool for accelerating AMP discovery and optimization.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
22.2%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.4%
13.1%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.7%
9.0%
4
PLOS Computational Biology
1863 papers in training set
Top 6%
6.3%
50% of probability mass above
5
Advanced Science
286 papers in training set
Top 0.8%
6.3%
6
Nature Communications
5641 papers in training set
Top 26%
5.6%
7
Bioinformatics Advances
203 papers in training set
Top 2%
3.3%
8
Scientific Reports
3612 papers in training set
Top 43%
2.4%
9
Cell Systems
201 papers in training set
Top 2%
2.4%
10
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 23%
2.2%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.8%
12
Bioinformatics
1204 papers in training set
Top 7%
1.7%
13
Cell Reports Methods
165 papers in training set
Top 2%
1.7%
14
iScience
1154 papers in training set
Top 24%
1.1%
15
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.9%
1.1%
16
PLOS ONE
5266 papers in training set
Top 54%
1.1%
17
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
18
Communications Chemistry
48 papers in training set
Top 1%
1.1%
19
Patterns
78 papers in training set
Top 3%
0.9%
20
Communications Biology
993 papers in training set
Top 29%
0.9%
21
PRX Life
42 papers in training set
Top 0.9%
0.9%
22
mAbs
32 papers in training set
Top 0.5%
0.6%
23
Protein Science
246 papers in training set
Top 4%
0.6%
24
BioData Mining
22 papers in training set
Top 1.0%
0.6%
25
Chemical Science
73 papers in training set
Top 2%
0.6%