SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information
Sun, M.; Wang, J.; Wan, S.
Show abstract
Antimicrobial resistance reduces the effectiveness of conventional antibiotics and has become a major global health threat, highlighting the need for new anti-infective agents. Antimicrobial peptides (AMPs), a diverse class of innate immune effectors with broad-spectrum antimicrobial activity, are promising candidates for combating drug-resistant infections. Identifying AMPs by wet-lab experiments, however, remains costly and time-consuming, creating a strong demand for computational identification methods. Our recently developed method, SAMP, captures region-specific residue distributions based on proportionalized split amino acid composition. However, SAMP might ignore key biochemical information and sequence order information. Here we present SAMP V2, a stacking ensemble learning framework based on biochemical and sequence-order information augmented split amino acid composition (BIA-SAAC), which extends SAMP by integrating pseudo-amino acid composition features with biochemical and sequence-order information into split peptide regions. Specifically, each peptide is divided into N-terminal, middle, and C-terminal regions, and pseudo amino acid composition is calculated within each region. Benchmarking tests on six independent test datasets, SAMP V2 outperformed multiple state-of-the-art models, including AMPpred-MFA and iAMP-Attenpred, in terms of accuracy, MCC, G-measure and F1-score. Given its high and robust performance, SAMP V2 could significantly accelerate the discovery of next-generation antimicrobial therapeutics for addressing the global threat of multidrug-resistant pathogens.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 95%
- InversePep: Diffusion-Driven Structure-Based Inverse Folding for Functional Peptides 94%
- Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy 93%
Similar papers in this journal
- High-PepBinder: A pLM-Guided Latent Diffusion Framework for Affinity-Aware Target-Specific Peptide Design 93%
- Graph Neural Networks Model Based on Atomic Hybridization for Predicting Drug Targets 93%
- Explainability methods from machine learning detect important drugs' atoms in drug-target interactions 93%
Similar papers in this journal
- A Multi-Property Optimizing Generative Adversarial Network for de novo Antimicrobial Peptide Design 95%
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 93%
- Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework 91%
Similar papers in this journal
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 93%
- Predicting Drug Protein Interaction using Quasi-Visual Question Answering System 92%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 92%
Similar papers in this journal
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 94%
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 93%
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.