Back

DeepAden: an explainable machine learning for substrate specificity prediction in nonribosomal peptide synthetases

Huang, J.; Ge, L.; Wu, Y.; Gao, Q.; Li, P.; Wu, J.; Zhang, H.; Qin, Z.

2025-05-26 bioinformatics
10.1101/2025.05.21.655435 bioRxiv
Show abstract

Microbial non-ribosomal peptides (NRPs) exhibit remarkable structural diversity and serve as valuable sources of lead compounds for clinical drug development. The biosynthesis of NRPs relies on non-ribosomal peptide synthetases (NRPSs), in which adenylation (A) domains play a pivotal role in defining the core structure by selectively recognizing and activating amino acid substrates. Accurately predicting the substrate specificities of A-domains is thus essential for understanding the core structural and biosynthetic logic of NRPs. Here, we present DeepAden, a two-stage deep learning framework. In the first stage, a graph attention network (GAT)-based model localizes 27-residue binding pockets within 6 [A] of bound substrates and convert these into pocket representations. In the second stage, pocket representations are then encoded alongside substrate information using pretrained language models, and aligned using contrastive learning. In addition, we introduce a SHapley Additive exPlanations (SHAP)-guided data augmentation strategy to mitigate class imbalance and improve robustness, particularly for nonproteinogenic substrates. DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions. DeepAden offers a powerful tool for precise pocket localization and robust substrate prediction, accelerating the discovery and characterization of novel NRP natural products for future work. The DeepAden web server is available at https://deepnp.site/.

Published in Nucleic Acids Research (predicted rank #10) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.