Back

PARAS: high-accuracy machine-learning of substrate specificities in nonribosomal peptide synthetases

Terlouw, B. R.; Huang, C.; Meijer, D.; Cediel-Becerra, J. D. D.; Rothe, M. L.; Jenner, M.; Zhou, S.; Zhang, Y.; Fage, C. D.; Tsunematsu, Y.; van Wezel, G. P.; Robinson, S. L.; Alberti, F. L.; Alkhalaf, L. M.; Chevrette, M. G.; Challis, G.; Medema, M. H.

2025-01-10 bioinformatics
10.1101/2025.01.08.631717 bioRxiv
Show abstract

Nonribosomal peptides are diverse natural products with important applications in medicine and agriculture. Bacterial and fungal genomes contain thousands of nonribosomal peptide biosynthetic gene clusters (BGCs) of unknown function, providing a promising resource for peptide discovery. Core structural features of such peptides can be inferred by predicting the substrate(s) of adenylation (A) domains in nonribosomal peptide synthetases (NRPSs). However, existing approaches to A domain prediction rely on limited datasets and often struggle with domains selecting large substrates or from less-studied taxa. Here, we systematically curate and computationally analyse 3,653 A domains and present two high-accuracy specificity predictors, PARAS and PARASECT. A type of A domain with unusually high L-tryptophan specificity was identified through the application of PARAS, and intact protein mass spectrometry to the corresponding NRPS showed it to direct the production of tryptopeptin-related metabolites in Streptomyces species. Together, these technologies will accelerate the characterisation of novel NRPSs and their metabolic products. PARAS and PARASECT are available at https://paras.bioinformatics.nl. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/631717v2_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@8e48aaorg.highwire.dtl.DTLVardef@144c78dorg.highwire.dtl.DTLVardef@890c11org.highwire.dtl.DTLVardef@1776de9_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in JACS Au · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.