Back

SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information

Sun, M.; Wang, J.; Wan, S.

2026-08-19 bioinformatics
10.64898/2026.08.12.744552 bioRxiv
Show abstract

Antimicrobial resistance reduces the effectiveness of conventional antibiotics and has become a major global health threat, highlighting the need for new anti-infective agents. Antimicrobial peptides (AMPs), a diverse class of innate immune effectors with broad-spectrum antimicrobial activity, are promising candidates for combating drug-resistant infections. Identifying AMPs by wet-lab experiments, however, remains costly and time-consuming, creating a strong demand for computational identification methods. Our recently developed method, SAMP, captures region-specific residue distributions based on proportionalized split amino acid composition. However, SAMP might ignore key biochemical information and sequence order information. Here we present SAMP V2, a stacking ensemble learning framework based on biochemical and sequence-order information augmented split amino acid composition (BIA-SAAC), which extends SAMP by integrating pseudo-amino acid composition features with biochemical and sequence-order information into split peptide regions. Specifically, each peptide is divided into N-terminal, middle, and C-terminal regions, and pseudo amino acid composition is calculated within each region. Benchmarking tests on six independent test datasets, SAMP V2 outperformed multiple state-of-the-art models, including AMPpred-MFA and iAMP-Attenpred, in terms of accuracy, MCC, G-measure and F1-score. Given its high and robust performance, SAMP V2 could significantly accelerate the discovery of next-generation antimicrobial therapeutics for addressing the global threat of multidrug-resistant pathogens.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.