Back

APDeeM: A machine Learning strategy towards Effective Peptide Vaccine Candidates Identification against Different Types of Viruses

Hossain, M. U.; Alom, M. R.; Hasan, S. S.; Sakib, M. N.; Suchi, M. A.; Sanjida, Z.; Rahman, A. B. Z. N.; Bhattacharjee, A.; Chowdhury, Z. M.; Ahammad, I.; Rahman, M. A.; Azad, S.; Salimullah, M.

2025-08-28 bioinformatics
10.1101/2025.08.25.671769 bioRxiv
Show abstract

Viral infections pose significant global health challenges, underscoring the urgent need for improved medications. Nevertheless, traditional medicinal approaches depend significantly on labor-intensive laboratory tests, which impede efficient identification and prolong vaccine development, particularly when screening a huge number of samples. To address these obstacles, we present a comprehensive Antiviral Peptide (AVP) Detection Dataset, comprising 14 unique features to improve the characterization of antiviral and non-antiviral peptides. Subsequently, we introduce the Antiviral Peptide detection enhanced by Ensemble Machine Learning (APDeeM) system. This advanced computational framework considerably reduces the time required for AVP detection by utilizing ensemble learning methodologies. The APDeeM system incorporates Gradient Boosting, Random Forest, K-Nearest Neighbors (KNN), and AdaBoost algorithms to facilitate the swift selection of AVP candidates without requiring urgent laboratory testing. Our proposed ensemble methodology showed superior performance, with an accuracy of 85.99%, F1 score of 87.60%, recall of 88.91%, and precision of 86.32%, exceeding the efficacy of all tested antiviral peptide prediction models in this research. The APDeeM approach signifies a substantial improvement over conventional detection techniques, expediting the identification of prospective vaccine candidates and facilitating the advancement of more effective antiviral peptides. The most promising AVP candidates may urge laboratory validation, optimize resources, and accelerate vaccine development.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.