Back

MetaPepticon: automated prediction of anticancer peptides from microbial genomes and metagenomes

Erözden, A. A.; Tavsanlı, N.; Demirel, G.; Sanlı, N. O.; Calıskan, M.; Arıkan, M.

2025-10-29 bioinformatics
10.1101/2025.10.28.685052 bioRxiv
Show abstract

Anticancer peptides (ACPs) are increasingly recognized as promising therapeutic candidates due to their ability to selectively target cancer cells. However, the systematic discovery of novel ACPs, particularly from high-throughput sequencing datasets, remains hindered by technical and methodological limitations. Current prediction frameworks require pre-extracted peptide sequences, involve manual preprocessing, and yield variable results, which restricts their applicability for large-scale, data-driven discovery. To address these limitations, we developed MetaPepticon, a modular, end-to-end pipeline for the discovery of candidate ACPs from diverse sequencing inputs, including raw genomic, metagenomic, transcriptomic, and metatranscriptomic reads, as well as assembled contigs and peptide sequences. MetaPepticon automates quality control, filtering, assembly, small open reading frame prediction, ACP classification using multiple predictive algorithms, and in silico toxicity filtering. By employing a consensus-based strategy and supporting heterogeneous data types, MetaPepticon facilitates scalable, reproducible, and high-confidence identification of candidate ACPs. Applied to 41,171 microbial genomes and 4,072,884 peptides, MetaPepticon identified 79,587 novel candidate ACPs, including 13,149 high-confidence, non-toxic peptides. By providing a standardized, automated framework for large-scale ACP discovery across various input types, MetaPepticon facilitates therapeutic peptide exploration and is freely available at: https://github.com/arikanlab/MetaPepticon

Published in PeerJ · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.