Fine-Grained Structural Classification of Biosynthetic Gene Cluster-Encoded Products
Porokhin, V.; Mevers, E.; van der Hooft, J. J. J.; Hassoun, S.
Show abstract
Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. Here, we introduce BGCat (BGC annotation tool), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles (PCPs) of gene cluster families (GCFs), associating each GCF with a probabilisitc distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 96%
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 95%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.