Back

Cost-effective hybrid long-short read assembly delineates alternative GC-rich Streptomyces chassis for natural product discovery

Heng, E.; Tan, L. L.; Tay, D. W. P.; Lim, Y. H.; Yang, L.-K.; Seow, D. C. S.; Leong, C. Y.; Ng, V.; Ng, S. B.; Kanagasundaram, Y.; Wong, F. T.; Koduru, L.

2022-12-06 genomics
10.1101/2022.12.05.519232 bioRxiv
Show abstract

With the advent of rapid automated in silico identification of biosynthetic gene clusters (BGCs), genomics presents vast opportunities to accelerate natural product (NP) discovery. However, prolific NP producers, Streptomyces, are exceptionally GC-rich (>80%) and highly repetitive within BGCs. These pose challenges in sequencing and high-quality genome assembly which are currently circumvented via intensive sequencing. Here, we outline a more cost-effective workflow using multiplex Illumina and Oxford Nanopore sequencing with hybrid long-short read assembly algorithms to generate high quality genomes. Our protocol involves subjecting long read-derived assemblies to up to 4 rounds of polishing with short reads to yield accurate BGC predictions. We successfully sequenced and assembled 8 GC-rich Streptomyces genomes whose lengths range from 7.1 to 12.1 Mb at an average N50 of 5.9 Mb. Taxonomic analysis revealed previous misrepresentation among these strains and allowed us to propose a potentially new species, Streptomyces sydneybrenneri. Further comprehensive characterization of their biosynthetic, pan-genomic and antibiotic resistance features especially for molecules derived from type I polyketide synthase (PKS) BGCs reflected their potential as NP chassis. Thus, the genome assemblies and insights presented here are envisioned to serve as gateway for the scientific community to expand their avenues in NP discovery. Graphic abstractSchematic of hybrid long- and short read assembly workflow for genome sequencing of GC-rich Streptomyces. Boxes shaded blue and grey correspond to experimental and in silico workflows, respectively. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=99 SRC="FIGDIR/small/519232v2_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@1395af0org.highwire.dtl.DTLVardef@81798forg.highwire.dtl.DTLVardef@539c92org.highwire.dtl.DTLVardef@14c366f_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIA cost-effective genome sequencing approach for GC-rich Streptomyces is presented C_LIO_LIHybrid assembly improves BGC annotation and identification C_LIO_LIA new species, Streptomyces sydneybrenneri, identified by taxonomic analysis C_LIO_LIGenomes of 8 Streptomyces species are reported and analysed in this study C_LI

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.