Back

Harnessing Uniform Design to Enhance AI-Driven Predictions of Physicochemical Properties of Short Peptides

Zhu, Z.; Liu, H.; Guo, Y.; Xu, M.; Li, X.; Zhou, H.; Wang, J.

2025-03-13 bioinformatics
10.1101/2025.03.10.642308 bioRxiv
Show abstract

Short peptides hold significant promise in drug discovery and materials science due to their biocompatibility, multifunctionality, and ease of synthesis. However, accurately predicting their physicochemical properties, a prerequisite for application development, remains a challenge. This study presents an innovative approach integrating uniform design (UD) with artificial intelligence (AI) to enhance prediction of key physicochemical properties, including aggregation propensity (AP), hydrophilicity (logP), and isoelectric point (pI). Using UD, we generate 31 distinct peptide datasets, with a consistent amino acid occupation fraction of 5% at each position, thereby creating unbiased training data for AI models. The performance of each AI model is rigorously evaluated using various testing schemes, and optimal sample sizes are determined for accurate prediction of each property. Additionally, Shapley Additive Explanations (SHAP) analysis identifies aromaticity, logP, net charge, and pI as the primary factors affecting peptide aggregation. This work provides comprehensive datasets on the physicochemical properties of all tetrapeptides, develops robust AI-based predictive models, and elucidates the relationships between key physicochemical characteristics and self-assembly behavior. By integrating experimental design, AI modeling, and peptide domain knowledge, our approach facilitates the discovery and optimization of functional peptides, offering new opportunities for peptide-based therapeutic applications.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Advanced Science
286 papers in training set
Top 0.3%
9.7%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.7%
7.8%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.9%
7.8%
4
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.2%
7.2%
5
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.3%
6.7%
6
Communications Chemistry
48 papers in training set
Top 0.1%
6.7%
7
Nature Communications
5641 papers in training set
Top 27%
5.4%
50% of probability mass above
8
The Journal of Physical Chemistry B
167 papers in training set
Top 0.7%
3.2%
9
iScience
1154 papers in training set
Top 7%
3.1%
10
Cell Reports Physical Science
19 papers in training set
Top 0.1%
2.8%
11
Chemical Science
73 papers in training set
Top 0.5%
2.8%
12
PLOS Computational Biology
1863 papers in training set
Top 11%
2.7%
13
Computers in Biology and Medicine
128 papers in training set
Top 2%
2.1%
14
Protein Science
246 papers in training set
Top 2%
1.9%
15
Scientific Reports
3612 papers in training set
Top 51%
1.9%
16
Small
78 papers in training set
Top 1.0%
1.7%
17
Biomacromolecules
29 papers in training set
Top 0.3%
1.5%
18
Analytical Chemistry
218 papers in training set
Top 2%
1.3%
19
Communications Biology
993 papers in training set
Top 22%
1.1%
20
ACS Omega
105 papers in training set
Top 2%
1.1%
21
ACS Central Science
71 papers in training set
Top 1%
1.1%
22
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
23
International Journal of Biological Macromolecules
76 papers in training set
Top 2%
1.0%
24
Molecular Therapy Nucleic Acids
39 papers in training set
Top 0.8%
1.0%
25
Nature Machine Intelligence
70 papers in training set
Top 2%
1.0%
26
Journal of Medicinal Chemistry
77 papers in training set
Top 0.9%
0.8%
27
Bioinformatics
1204 papers in training set
Top 9%
0.8%
28
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 45%
0.6%
29
PLOS ONE
5266 papers in training set
Top 65%
0.6%