Back

GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13 Inducing Peptides

Kurata, H.; Tsuruta, H.; Shigetomi, S.; HARUN-OR-ROSHID, M.; Maeda, K.

2025-12-10 bioinformatics
10.64898/2025.12.07.692883 bioRxiv
Show abstract

The identification of interleukin-6 (IL-6) and interleukin-13 (IL-13) inducing peptides is crucial for accelerating drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening methods remain time-consuming and costly. To address these limitations, various machine learning and deep learning models have been developed, yet their performance is still constrained by the limited availability of experimentally validated data. In this study, we propose a generative AI-driven data augmentation (GDA) framework and predictor, GDA-Pred, to improve the prediction performance of state-of-the-art (SOTA) classifiers for identifying IL-6 and IL-13 inducing peptides. GDA expands the training dataset by generating novel peptide sequences using three types of generative AI models: generative adversarial networks (GANs), diffusion models (DMs), and variational autoencoders (VAEs). The GDA framework is defined by four key parameters: the type of generative model, the sequence identity cutoff, the probability threshold (PT) for selecting generated peptides, and the augmentation ratio (AR) between generated and real peptides. Since optimizing these parameters is challenging with small datasets, we adopt a case study-oriented proof-of-concept approach using a moderately sized dataset of anti-inflammatory peptides (AIPs) to derive interpretable optimal settings. The performance of the optimized GDA was evaluated using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out test on benchmark datasets. The optimized GDA was then applied to SOTA classifiers, collectively termed GDA-Pred, to identify IL-6 and IL-13 inducing peptides, both of which are limited by small dataset sizes. GDA-Pred substantially improved the prediction performance for both cytokine-inducing peptide tasks, demonstrating the feasibility of GDA-Pred as a robust and generalizable framework.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.