Nerpa 2: linking biosynthetic gene clusters to nonribosomal peptide structures
Olkhovskii, I.; Kushnareva, A.; Popov, P.; Tagirdzhanov, A.; Gurevich, A.
Show abstract
MotivationNonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways. ResultsWe present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC-product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry. Availability and implementationNerpa 2 is freely available at https://github.com/gurevichlab/nerpa.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- IPANEMAP: Integrative Probing Analysis of Nucleic Acids Empowered by Multiple Accessibility Profiles 93%
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 93%
- The ModelSEED Database for the integration of metabolic annotations and the reconstruction, comparison, and analysis of metabolic models for plants, fungi, and microbes 92%
Similar papers in this journal
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 94%
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 93%
- DepoScope: accurate phage depolymerase annotation and domain delineation using large language models 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.