GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13 Inducing Peptides
Kurata, H.; Tsuruta, H.; Shigetomi, S.; HARUN-OR-ROSHID, M.; Maeda, K.
Show abstract
The identification of interleukin-6 (IL-6) and interleukin-13 (IL-13) inducing peptides is crucial for accelerating drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening methods remain time-consuming and costly. To address these limitations, various machine learning and deep learning models have been developed, yet their performance is still constrained by the limited availability of experimentally validated data. In this study, we propose a generative AI-driven data augmentation (GDA) framework and predictor, GDA-Pred, to improve the prediction performance of state-of-the-art (SOTA) classifiers for identifying IL-6 and IL-13 inducing peptides. GDA expands the training dataset by generating novel peptide sequences using three types of generative AI models: generative adversarial networks (GANs), diffusion models (DMs), and variational autoencoders (VAEs). The GDA framework is defined by four key parameters: the type of generative model, the sequence identity cutoff, the probability threshold (PT) for selecting generated peptides, and the augmentation ratio (AR) between generated and real peptides. Since optimizing these parameters is challenging with small datasets, we adopt a case study-oriented proof-of-concept approach using a moderately sized dataset of anti-inflammatory peptides (AIPs) to derive interpretable optimal settings. The performance of the optimized GDA was evaluated using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out test on benchmark datasets. The optimized GDA was then applied to SOTA classifiers, collectively termed GDA-Pred, to identify IL-6 and IL-13 inducing peptides, both of which are limited by small dataset sizes. GDA-Pred substantially improved the prediction performance for both cytokine-inducing peptide tasks, demonstrating the feasibility of GDA-Pred as a robust and generalizable framework.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 95%
- DeepTAP: an RNN-based method of TAP-binding peptide prediction in the selection of tumor neoantigens 94%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 93%
Similar papers in this journal
- Personalized deep learning of individual immunopeptidomes to identify neoantigens for cancer vaccines 94%
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 94%
- A deep learning framework for high-throughput mechanism-driven phenotype compound screening 93%
Similar papers in this journal
Similar papers in this journal
- A Multi-Property Optimizing Generative Adversarial Network for de novo Antimicrobial Peptide Design 97%
- Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework 93%
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.