Back

Machine Learning Uncovers the Transcriptional Regulatory Network for the Production Host Streptomyces albidoflavus J1074

Jönsson, M.; Sigrist, R.; Semenov Petrov, M.; Marcussen, N.; Gren, T.; Palsson, B. O.; Yang, L.; Özdemir, E.

2024-01-10 genomics
10.1101/2024.01.09.574332 bioRxiv
Show abstract

Streptomyces albidoflavus is a popular and genetically tractable platform strain used for natural product discovery and production via the expression of heterologous biosynthetic gene clusters (BGCs). However, its transcriptional regulatory network (TRN) and its impact on secondary metabolism is poorly understood. Here we characterized its TRN by applying an independent component analysis to a compendium of 218 high quality RNA-seq transcriptomes from both in-house and public sources spanning 88 unique growth conditions. We obtained 78 independently modulated sets of genes (iModulons) that quantitatively describe the TRN and its activity state across diverse conditions. Through analyses of condition-dependent TRN activity states, we (i) describe how the TRN adapts to different growth conditions, (ii) conduct a cross-species iModulon comparison, uncovering shared features and unique characteristics of the TRN across lineages, (iii) detail the transcriptional activation of several endogenous BGCs, including surugamide, minimycin and paulomycin, and (iv) infer potential functions of 40% of the uncharacterized genes in the S. albidoflavus genome. Our findings provide a comprehensive and quantitative understanding of the TRN of S. albidoflavus, providing a knowledge base for further exploration and experimental validation. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=198 SRC="FIGDIR/small/574332v3_ufig1.gif" ALT="Figure 1"> View larger version (44K): org.highwire.dtl.DTLVardef@1a7d0aorg.highwire.dtl.DTLVardef@10750eborg.highwire.dtl.DTLVardef@1518644org.highwire.dtl.DTLVardef@145f5b4_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.