Predicting microbial transcriptome using genome sequence
SHAO, B.; YAN, Y.; FU, G.
Show abstract
We present TXpredict, a transformer-based framework for predicting microbial transcriptomes using annotated genome sequences. By leveraging information learned from a large protein language model, TXpredict achieves an average Spearman correlation of 0.53 and 0.62 in predicting gene expression for new bacterial and fungal genomes. We further extend this framework to predict transcriptomes for 2, 685 additional microbial genomes spanning 1, 744 genera, 82% of which remain uncharacterized at the transcriptional level. Our analysis highlights conserved and divergent transcriptional programs across understudied genera, providing a powerful resource for uncovering microbial adaptation strategies and metabolic potential across the tree of life. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=103 SRC="FIGDIR/small/630741v4_ufig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@d1cad1org.highwire.dtl.DTLVardef@15a9206org.highwire.dtl.DTLVardef@128e65borg.highwire.dtl.DTLVardef@2b7ca7_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- LinearTurboFold: Linear-Time Global Prediction of Conserved Structures for RNA Homologs with Applications to SARS-CoV-2 95%
- Contrasting patterns of microbial dominance in the Arabidopsis thaliana phyllosphere 95%
- Revealing 29 sets of independently modulated genes in Staphylococcus aureus, their regulators and role in key physiological responses 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.