Back

Pathogen context reshapes antimicrobial peptide generation

You, S.; Zhang, C.; Han, Y.; Jiang, Q.; Guo, X.; Li, M.; Su, Y.; Dong, X.; Yang, M.; Lu, H.

2026-07-03 bioinformatics
10.64898/2026.07.01.735178 bioRxiv
Show abstract

Antimicrobial peptide discovery is constrained less by the number of molecules that can be generated than by the choice of which few should be tested against a defined pathogen. Most peptide generators produce broadly antimicrobial-like sequences and leave target specificity to downstream filters. Here we show that pathogen context can be introduced during generation. AMPHORA conditions a peptide-native generator on target class, pathogen genome features and strain-description text. Matched, ablated and shuffled controls showed that aligned pathogen inputs redirected generated libraries beyond coarse activity labels, whereas global shuffling weakened this effect. Same-noise counterfactuals showed that strain descriptions drove larger sequence changes, whereas genome features more strongly affected predicted structural properties. Species-level analyses revealed target-dependent enrichment. Matched bacterial inputs also shifted APEX-predicted activity rankings relative to class-only generation. The resulting libraries remained diverse, largely non-memorizing and compatible with predicted peptide-like structural features. Together, these results establish pathogen-context conditioning as a new paradigm for computational library reshaping in antimicrobial peptide generation.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
18.2%
2
Nature Communications
5641 papers in training set
Top 18%
9.6%
3
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 5%
7.8%
4
PLOS Computational Biology
1863 papers in training set
Top 5%
7.1%
5
Nature Machine Intelligence
70 papers in training set
Top 0.8%
3.5%
6
Advanced Science
286 papers in training set
Top 2%
3.2%
7
Nature Biotechnology
172 papers in training set
Top 1%
3.2%
50% of probability mass above
8
eLife
5828 papers in training set
Top 35%
3.2%
9
Molecular Systems Biology
162 papers in training set
Top 0.9%
2.6%
10
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.4%
11
iScience
1154 papers in training set
Top 14%
2.1%
12
Scientific Reports
3612 papers in training set
Top 48%
2.1%
13
Genome Biology
637 papers in training set
Top 5%
1.9%
14
Science
477 papers in training set
Top 5%
1.7%
15
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
16
Nature Biomedical Engineering
47 papers in training set
Top 0.8%
1.7%
17
Nature Chemical Biology
119 papers in training set
Top 2%
1.7%
18
Cell
431 papers in training set
Top 7%
1.5%
19
Science Advances
1243 papers in training set
Top 22%
1.5%
20
Cell Genomics
172 papers in training set
Top 3%
1.1%
21
Communications Chemistry
48 papers in training set
Top 1%
1.1%
22
Cell Reports
1498 papers in training set
Top 25%
1.0%
23
PLOS ONE
5266 papers in training set
Top 58%
1.0%
24
Nature
645 papers in training set
Top 10%
1.0%
25
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
26
Nature Methods
385 papers in training set
Top 6%
0.8%
27
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
28
Chemical Science
73 papers in training set
Top 2%
0.8%
29
ACS Synthetic Biology
287 papers in training set
Top 3%
0.6%
30
ACS Central Science
71 papers in training set
Top 2%
0.6%