Back

Sequence-Based Therapeutic Peptide Classification with Augmented Negative Sampling

Ellerbrock, R.; Valentini, A.; Paul, A. C.; Mukhopadhyay, S.; Perelshtein, M. R.

2026-06-11 bioinformatics
10.64898/2026.06.07.730473 bioRxiv
Show abstract

Therapeutic peptides offer high target specificity, low toxicity, and the ability to modulate protein-protein interactions, yet experimental functional characterization remains costly and slow. Computational prediction of therapeutic function directly from sequence could accelerate peptide screening and enable generative design pipelines, but requires reliable discrimination between therapeutic and non-therapeutic peptides. Existing multi-label predictors cover few functions, rely on limited datasets, and exhibit high False Positive Rates (FPRs), limiting their practical utility. We present a lightweight CNN classifier trained on the most comprehensive therapeutic peptide database to date (54,655 peptides, 48 functional categories). A key contribution is a statistically motivated negative sampling strategy using Markov models to generate diverse synthetic decoys at multiple difficulty levels. When evaluated on this controlled decoy benchmark, the FPR is reduced from over 60% for previous models to 2.1% for our approach. On positive therapeutic samples, our fine-tuned five-model ensemble achieves 79.9% Micro F1 and 54.6% Macro F1 while requiring only amino acid sequences as inputs. Analysis using a sparse L1-constrained variant of our model shows that convolutional filters capture conserved functional motifs and statistically improbable non-therapeutic patterns, with downstream layers combining these signals, providing mechanistic evidence that the network learns biologically meaningful structure. On an external generalization benchmark derived from TPpred-LE, our model achieves 55.3% Micro F1 and 38.6% Macro F1 on the 12 shared labels, close to the benchmark-specific baseline (57.9%/38.1%), while retaining substantially broader therapeutic label coverage. Code and models will be made available at https://github.com/terra-quantum-public/tq-therapep-ai.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
26.8%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.2%
18.7%
3
mAbs
32 papers in training set
Top 0.1%
5.2%
50% of probability mass above
4
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.9%
5
Nature Communications
5641 papers in training set
Top 31%
4.4%
6
Bioinformatics
1204 papers in training set
Top 5%
3.3%
7
Cell Systems
201 papers in training set
Top 2%
2.8%
8
Communications Chemistry
48 papers in training set
Top 0.3%
2.4%
9
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 26%
1.9%
10
Chemical Science
73 papers in training set
Top 0.9%
1.7%
11
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
12
Nature Biotechnology
172 papers in training set
Top 2%
1.7%
13
Nature Methods
385 papers in training set
Top 4%
1.7%
14
Patterns
78 papers in training set
Top 2%
1.1%
15
Scientific Reports
3612 papers in training set
Top 64%
1.1%
16
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.9%
1.1%
17
Advanced Science
286 papers in training set
Top 7%
1.1%
18
Nature Biomedical Engineering
47 papers in training set
Top 1%
1.0%
19
Cell Reports Methods
165 papers in training set
Top 3%
1.0%
20
Journal of Cheminformatics
29 papers in training set
Top 0.6%
1.0%
21
ACS Central Science
71 papers in training set
Top 1%
0.9%
22
Protein Science
246 papers in training set
Top 3%
0.9%
23
PRX Life
42 papers in training set
Top 1%
0.6%
24
eLife
5828 papers in training set
Top 68%
0.6%
25
Molecular Systems Biology
162 papers in training set
Top 4%
0.6%