NPannotator: a genome- and chemistry- constrained automation for type I polyketide synthase pathway elucidation
Chainani, Y.; Cornman, A.; Hwang, Y.
Show abstract
Natural products (NPs) are structurally diverse bioactive compounds whose biosynthesis is encoded within biosynthetic gene clusters (BGCs). Although databases such as the Minimum Information about a Biosynthetic Gene Cluster (MiBIG) repository now catalog thousands of experimentally validated NP structures, the full biosynthetic pathway connecting individual domain sequences to specific chemical features on final NP structures remains largely unannotated. This gap is especially pronounced for type I polyketide synthases (PKSs). These are modular assembly lines in which multiple enzymatic domains work in concert to condense acyl-CoA building blocks into complex polyketide scaffolds. Within these systems, acyltransferase (AT) domains govern which starter and extender units are incorporated at each elongation step, yet the substrate specificities of AT domains are known for only a fraction of cataloged clusters. Moreover, the catalytic order of genes encoding PKS modules is not immediately apparent from existing database entries, leaving the correct module ordering for observed product structures uncertain. Here, we present NPannotator, an automated, genomic context-aware cheminformatics pipeline that infers both the catalytic ordering of PKS domains and the substrate specificities of a given PKSs AT domains. NPannotator loads a precomputed database of synthetically generated polyketide backbones, iteratively replaces default malonyl-CoA substrates with candidate starter and extender units via SMARTS-based substructure matching against the target NP, and selects the arrangement that maximizes chemical similarity. When benchmarked on the type I PKSs annotated within the expert-reviewed ClusterCAD dataset, NPannotator recovered 62.0% of both correct gene orderings and AT substrate annotations, and achieved 80.0% accuracy on gene ordering alone. By bridging gene-level architecture with chemical outcomes, NPannotator represents a step toward systematically decoding how protein sequence and genomic organization encode chemical structure in the world of natural products.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing pathways for bioproducing complex chemicals by combining tools for pathway extraction and ranking 96%
- Galaxy-SynBioCAD: Automated Pipeline for Synthetic Biology Design and Engineering 95%
- Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production 94%
Similar papers in this journal
- Integration of diverse bioactivity data into the Chemical Checker compound universe 93%
- OCTAD: an open workplace for virtually screening therapeutics targeting precise cancer patient groups using gene expression features 92%
- Measuring carbohydrate recognition profile of lectins on live cells using liquid glycan array (LiGA) 90%
Similar papers in this journal
Similar papers in this journal
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 94%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 94%
- Transfer learning enables prediction of CYP2D6 haplotype function 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.