DeepSeMS: a large language model reveals hidden biosynthetic potential of the global ocean microbiome
Xu, T.; Yang, Y.; Zhu, R.; Lin, W.; Li, J.; Zheng, Y.; Zhang, P.; Zhang, G.; Zhao, G.-P.; Jiao, N.
Show abstract
Microbial biosynthetic diversity holds immense potential for discovering natural products with therapeutic applications, yet a substantial quantity of natural products derived from uncultivated microorganisms remains uncharacterized. The intricate nature of biosynthetic enzymes poses a major challenge in accurately predicting the chemical structures of secondary metabolites solely based on genome sequences using current rule-based methods. Here, we present DeepSeMS, a large language model designed to predict the chemical structures of secondary metabolites from various microbial biosynthetic gene clusters. Built on the Transformer architecture, DeepSeMS innovatively identifies sequence features using functional domains of biosynthetic enzymes, and incorporates feature-aligned chemical structure enumeration for training data augmentation. External evaluation results show that DeepSeMS predicts more accurate chemical structures of secondary metabolites with a Tanimoto coefficient up to 0.6 compared with the ground truth, significantly outperforming antiSMASH and PRISM with coefficients of only 0.14 and 0.45 respectively. Moreover, DeepSeMS successfully predicted secondary metabolites for 96.60% of cryptic biosynthetic gene clusters, surpassing existing methods with success rates less than 50%. Leveraging DeepSeMS, we characterized over 65,000 novel secondary metabolites from the global ocean microbiome with previously undocumented structural types, ecological distribution, and biomedical applications especially antibiotics. A login-free and user-friendly web server for DeepSeMS (https://biochemai.cstspace.cn/deepsems/) has been launched, featuring an integrated global ocean microbial secondary metabolites repository to expedite the discovery of novel natural products. Collectively, this study underscores the great capacity of a large language model-driven method in revealing hidden biosynthetic potential of the global ocean microbiome.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An artificial intelligence accelerated virtual screening platform for drug discovery 94%
- Biosynthetic Enzyme-guided Disease Correlation Connects Gut Microbial Metabolites Sulfonolipids to Inflammatory Bowel Disease Involving TLR4 Signaling 94%
- AGILE Platform: A Deep Learning-Powered Approach to Accelerate LNP Development for mRNA Delivery 94%
Similar papers in this journal
- ProT-Diff: A Modularized and Efficient Approach to De Novo Generation of Antimicrobial Peptide Sequences through Integration of Protein Language Model and Diffusion Model 96%
- Transfer Learning and Permutation-Invariance improving Predicting Genome-wide, Cell-Specific and Directional Interventions Effects of Complex Systems 95%
- A Multi-Property Optimizing Generative Adversarial Network for de novo Antimicrobial Peptide Design 93%
Similar papers in this journal
- The first archaeal PET-degrading enzyme belongs to the feruloyl-esterase family 92%
- EVOSYNTH: Enabling Multi-Target Drug Discovery through Latent Evolutionary Optimization and Synthesis-Aware Prioritization 92%
- Metabolic engineering and late-stage functionalization expand the chemical space of the antimalarial premarineosin A 91%
Similar papers in this journal
- Large Scale Cell Painting Guided Compound Selection Reveals Activity Cliffs and Functional Relationships 95%
- Exploration of natural red-shifted rhodopsins using a machine learning-based Bayesian experimental design 94%
- Long-read metagenomics of soil communities reveals phylum-specific secondary metabolite dynamics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.