A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides
Abhigyan, R.; Sood, V.; Arora, P.; Kaur, B.
Show abstract
Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative-evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. HighlightsO_LIProposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. C_LIO_LIPerformed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. C_LIO_LICombined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. C_LIO_LIDemonstrated the applicability of the proposed framework across multiple bioactive peptide datasets. C_LI
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 93%
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 92%
- DeepSP: Deep Learning-Based Spatial Properties to Predict Monoclonal Antibody Stability 92%
Similar papers in this journal
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 94%
- InversePep: Diffusion-Driven Structure-Based Inverse Folding for Functional Peptides 93%
- ProDualNet: Dual-Target Protein Sequence Design Method Based on Protein Language Model and Structure Model 93%
Similar papers in this journal
- Peptide-aware chemical language model successfully predicts membrane diffusion of cyclic peptides 93%
- Pred-AHCP: Robust feature selection enabled Sequence Specific Prediction of Anti-Hepatitis C Peptides via Machine Learning 93%
- WITHDRAWN: The Use of DeepQSAR Models for The Discovery of Peptides with Enhanced Antimicrobial and Antibiofilm Potential. 93%
Similar papers in this journal
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 93%
- Personalized deep learning of individual immunopeptidomes to identify neoantigens for cancer vaccines 91%
- A deep learning framework for high-throughput mechanism-driven phenotype compound screening 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.