Back

Expansion of novel biosynthetic gene clusters from diverse environments using SanntiS

Sanchez Fragoso, J. S.; Rogers, J. D.; Rogers, A. B.; Nassar, M.; McEntyre, J.; Welch, M.; Hollfelder, F.; Finn, R. D.

2023-05-24 bioinformatics
10.1101/2023.05.23.540769 bioRxiv
Show abstract

Natural products biosynthesised by microbes are an important component of the pharmacopeia with a vast array of biomedical and industrial applications, in addition to their key role in mediating many ecological interactions. One approach for the discovery of these metabolites is the identification of biosynthetic gene clusters (BGCs), genomic units which encode the molecular machinery required for producing the natural product. Genome mining has revolutionised the discovery of BGCs, yet metagenomic assemblies represent a largely untapped source of natural products. The imbalanced distribution of BGC classes in existing databases restricts the generalisation of detection patterns and limits the ability of mining methods to recognise a broader spectrum of BGCs. This problem is further intensified in metagenomic datasets, where BGC genes may be incomplete. This work presents SanntiS, a new machine learning-based tool for identifying BGCs. SanntiS achieved high precision and recall in both genomic and metagenomic datasets, effectively capturing a broad range of BGCs. Application of SanntiS to MGnify metagenomic assemblies led to a resource containing 1.9 million BGC predictions with associated contextual data from diverse biomes and demonstrates a significant fraction of novelty compared to equivalent isolate genomes datasets. Subsequent experimental validation of a novel antimicrobial peptide detected solely by SanntiS, further demonstrates the potential of this approach for uncovering novel bioactive compounds.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.