Back

Pangenome mining of the Streptomyces genus redefines their biosynthetic potential

Mohie, O. S.; Jorgensen, T. S.; Booth, T. J.; Charusanti, P.; Phaneuf, P. V.; Weber, T.; Palsson, B. O.

2024-02-22 systems biology
10.1101/2024.02.20.581055 bioRxiv
Show abstract

BackgroundStreptomyces is a highly diverse genus known for the production of secondary or specialized metabolites with a wide range of applications in the medical and agricultural industries. Several thousand complete or nearly-complete Streptomyces genome sequences are now available, affording the opportunity to deeply investigate the biosynthetic potential within these organisms and to advance natural product discovery initiatives. ResultWe performed pangenome analysis on 2,371 Streptomyces genomes, including approximately 1,200 complete assemblies. Employing a data-driven approach based on genome similarities, the Streptomyces genus was classified into 7 primary and 42 secondary MASH-clusters, forming the basis for a comprehensive pangenome mining. A refined workflow for grouping biosynthetic gene clusters (BGCs) redefined their diversity across different MASH-clusters. This workflow also reassigned 2,729 known BGC families to only 440 families, a reduction caused by inaccuracies in BGC boundary detections. When the genomic location of BGCs is included in the analysis, a conserved genomic structure (synteny) among BGCs becomes apparent within species and MASH-clusters. This synteny suggests that vertical inheritance is a major factor in the acquisition of new BGCs. ConclusionOur analysis of a genomic dataset at a scale of thousands of genomes refined predictions of BGC diversity using MASH-clusters as a basis for pangenome analysis. The observed conservation in the order of BGCs genomic locations showed that the BGCs are vertically inherited. The presented workflow and the in-depth analysis pave the way for large-scale pangenome investigations and enhance our understanding of the biosynthetic potential of the Streptomyces genus.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.