CORSID enables de novo identification of transcription regulatory sequences and genes in coronaviruses
Zhang, C.; Sashittal, P.; El-Kebir, M.
Show abstract
Genes in coronaviruses are preceded by transcription regulatory sequences (TRSs), which play a critical role in gene expression mediated by the viral RNA-dependent RNA-polymerase via the process of discontinuous transcription. In addition to being crucial for our understanding of the regulation and expression of coronavirus genes, we demonstrate for the first time how TRSs can be leveraged to identify gene locations in the coronavirus genome. To that end, we formulate the TRS AND GO_SCPLOWENEC_SCPLOW IO_SCPLOWDENTIFICATIONC_SCPLOW (TRS-GO_SCPLOWENEC_SCPLOW-ID) problem of simultaneously identifying TRS sites and gene locations in unannotated coronavirus genomes. We introduce CORSID (CORe Sequence IDentifier), a computational tool to solve this problem. We also present CORSID-A, which solves a constrained version of the TRS-GO_SCPLOWENEC_SCPLOW-ID problem, the TRS IO_SCPLOWDENTIFICATIONC_SCPLOW (TRS-ID) problem, identifying TRS sites in a coronavirus genome with specified gene annotations. We show that CORSID-A outperforms existing motif-based methods in identifying TRS sites in coronaviruses and that CORSID outperforms state-of-the-art gene finding methods in finding genes in coronavirus genomes. We demonstrate that CORSID enables de novo identification of TRS sites and genes in previously unannotated coronaviruses. CORSID is the first method to perform accurate and simultaneous identification of TRS sites and genes in coronavirus genomes without the use of any prior information.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Single-cell tumor phylogeny inference with copy-number constrained mutation losses 94%
- Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics 93%
- RepairSig: Deconvolution of DNA damage and repaircontributions to the mutational landscape of cancer 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.