Back

AI-driven discovery and optimization of antimicrobial peptides from extreme environments on global scale

Kang, Z.; Zhang, H.; Zhou, Q.; Liu, J.; Zhou, K.; Chen, P.; Liu, B.-F.; Ning, K.

2025-11-26 bioinformatics
10.1101/2025.11.13.688364 bioRxiv
Show abstract

The escalating crisis of global antimicrobial resistance (AMR) necessitates the discovery of novel antibiotics. Antimicrobial peptides (AMPs), particularly those from under-explored extreme environments, represent a promising therapeutic class. Here, we introduce SEGMA (Structure-aware Extremophile Genome Mining for Antimicrobial peptides), a computational framework that integrates structure information to systematically mine AMPs from extremophile genomes on a global scale. By analyzing 60,461 extremophile metagenome-assembled genomes (MAGs) from diverse habitats, SEGMA identified 3,298 novel AMPs (termed "extremocins"), which exhibit unique amino acid profiles and physicochemical properties. Leveraging a beam search-guided optimization strategy, we further enhanced selected extremocins to achieve broad-spectrum antimicrobial activity. Experimental validation confirmed potent in vitro efficacy against clinically relevant pathogens. This study underscores the value of structure-aware mining and extremophile microbiomes in expanding the antibiotic arsenal against AMR. HighlightsO_LISEGMA, a structure-aware deep learning framework, mines 3,298 novel antimicrobial peptides (extremocins) from 60,461 extremophile genomes on global scale. C_LIO_LIExtremocins exhibit unique sequence features, and expand known antibiotic space, few of which shows homology to existing AMP databases. C_LIO_LIA beam search-guided optimization strategy enhanced selected extremocins to achieve broad-spectrum activity against clinically relevant pathogens. C_LIO_LIExperimental validation confirmed that candidate extremocins exhibit potent in vitro and in vivo antimicrobial activity, highlighting their therapeutic potential. C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.