How Does Sampling Affect the AI Prediction Accuracy of Peptides' Physicochemical Properties?
Yan, M.; Abuduhebaier, A.; Zhou, H.; Wang, J.
10.1101/2025.01.29.635451 bioRxivShow abstract
Accurate AI prediction of peptide physicochemical properties is essential for advancing peptide-based biomedicine, biotechnology, and bioengineering. However, the performance of predictive AI models is significantly affected by the representativeness of the training data, which depends on the sample size and sampling methods employed. This study addresses the challenge of determining the optimal sample size and sampling methods to enhance the predictive accuracy and generalization capacity of AI models for estimating the aggregation propensity, hydrophilicity, and isoelectric point of tetrapeptides. Four sampling methods were evaluated: Latin Hypercube Sampling (LHS), Uniform Design Sampling (UDS), Simple Random Sampling (SRS), and Probability-Proportional-to-Size Sampling (PPS), across sample sizes ranging from 100 to 20,000. A sample size of approximately 12,000 (7.5% of the total tetrapeptide dataset) marks a key threshold for stable and consistent model performance. This study provides valuable insights into the interplay between sample size, sampling strategies, and model performance, offering a foundational framework for optimizing data collection and AI model training for the prediction of peptides physicochemical properties, especially for prediction in the complete sequence space of longer peptides with more than four amino acids.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accurate and Rapid Prediction of Protein pKa: Protein Language Models Reveal the Sequence-pKa Relationship 95%
- BioStructNet: Structure-Based Network with Transfer Learning for Predicting Biocatalyst Functions 95%
- ANUBI: A Platform for Affinity Optimization of Proteins and Peptides in Drug Design 95%
Similar papers in this journal
- DeepSP: Deep Learning-Based Spatial Properties to Predict Monoclonal Antibody Stability 96%
- DeepSCM: an efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity 95%
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 94%
Similar papers in this journal
Similar papers in this journal
- Computational Modeling of Stapled Coiled-Coil Inhibitors Against Bcr-Abl: Toward a Treatment Strategy for CML 96%
- The origin of secondary structure transitions and peptide self-assembly propensity in trifluoroethanol-water mixtures 95%
- Predicting the Activities of Drug Excipients on Biological Targets using One-Shot Learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.