TSdb: A Curated Database for Terpene Synthases and Their Application in Natural Product Mining
Gu, T.; Ming, D.
Show abstract
Terpenoids constitute natures most chemically diverse metabolite family with vital pharmaceutical and industrial applications, yet existing databases lack systematic integration of precursor metabolic enzymes (HMGR, DXS) and mechanistic insights into terpene diversification. To bridge this gap, we developed the Terpene Synthase Database (TSDB), distinguishing itself through three key innovations: (1) comprehensive integration of MVA/MEP pathway enzymes with downstream terpenoid synthases, (2) enhanced functional annotation via InterProScan domain mapping and phylogenetics to decode catalytic plasticity, and (3) unprecedented taxonomic breadth spanning 456,142 non-redundant sequences across 30,491 taxa. By consolidating data from BRENDA, UniProt, TeroKit, and ocean gene clusters through rigorous BLASTp deduplication (95% identity cutoff), TSDB reveals 3,499 Gene Ontology terms highlighting core functions like isoprenoid biosynthesis (GO:0019288) and metalloenzyme catalysis. Validation against MIBiG gene clusters (e.g., BGC0001324) demonstrates precise identification of terpene cyclases, P450 monooxygenases, and prenyltransferases with residue-level active site annotations. As the first resource connecting precursor metabolism to structural diversity, TSDB enables accurate gene-enzyme-product prediction for enzyme engineering and natural product discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis of secondary metabolite gene clusters and chitin biosynthesis pathways of Monascus purpureus with high production of pigment and citrinin based on whole-genome sequencing 96%
- Metabolomics and Transcriptomics unravel the mechanism of browning resistance in Agaricus bisporus 96%
- In Silico Identification, Characterization and Diversity Analysis of RNAi Genes and their Associated Regulatory Elements in Sweet Orange (Citrus sinensis L) 95%
Similar papers in this journal
- CNSA: a data repository for archiving omics data 95%
- GenDiS3 database: census on prevalence of protein domain superfamilies of known structure in the entire sequence database 95%
- CitrusKB: A Comprehensive Knowledge Base for Transcriptome and Interactome of Citrus spp. Infected by Xanthomonas citri subsp. citri at Different Infection Stages 94%
Similar papers in this journal
- FishExp: a comprehensive database and analysis platform for gene expression and alternative splicing of fish species 94%
- AutoVEM2: a flexible automated tool to analyze candidate key mutations and epidemic trends for virus 93%
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 93%
Similar papers in this journal
- Photosynthetic protein classification using genome neighborhood-based machine learning feature 94%
- Comparative Analysis, Diversification and Functional Validation of Plant Nucleotide-Binding Site Domain Genes 94%
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.