Analysis of next- and third-generation RNA-Seq data reveals the structures of alternative transcription units in bacterial genomes
Wang, Q.; Liu, Z.; Yan, B.; Chou, W.-C.; Ettwiller, L. M.; Ma, Q.; Liu, B.
Show abstract
Alternative transcription units (ATUs) are dynamically encoded under different conditions or environmental stimuli in bacterial genomes, and genome-scale identification of ATUs is essential for studying the emergence of human diseases caused by bacterial organisms. However, it is unrealistic to identify all ATUs using experimental techniques, due to the complexity and dynamic nature of ATUs. Here we present the first-of-its-kind computational framework, named SeqATU, for genome-scale ATU prediction based on next-generation RNA-Seq data. The framework utilizes a convex quadratic programming model to seek an optimum expression combination of all of the to-be-identified ATUs. The predicted ATUs in E. coli reached a precision of 0.77/0.74 and a recall of 0.75/0.76 in the two RNA-Sequencing datasets compared with the benchmarked ATUs from third-generation RNA-Seq data. We believe that the ATUs identified by SeqATU can provide fundamental knowledge to guide the reconstruction of transcriptional regulatory networks in bacterial genomes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Species-specific design of artificial promoters by transfer-learning based generative deep-learning model 96%
- Best practices for perturbation MPRA--a computational evaluation framework of sequence design strategies 95%
- iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning 94%
Similar papers in this journal
Similar papers in this journal
- TAMC: A deep-learning approach to predict motif-centric transcriptional factor binding activity based on ATAC-seq profile 93%
- RNANetMotif: identifying sequence-structure RNA network motifs in RNA-protein binding sites 93%
- Base-resolution prediction of transcription factor binding signals by a deep learning framework 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.