CFC-seq: identification of full-length capped RNAs unveil enhancer-derived transcription
Yip, C. W.; Parr, C.; Takahashi, H.; Yasuzawa, K.; Valentine, M.; Nishiyori-Sueki, H.; Ugolini, C.; Ranzani, V.; Murata, M.; Kato, M.; Kang, W.; Yip, W. H.; Shibayama, Y.; Sim, A. D.; Chen, Y.; Shu, X.; Moody, J. D.; Umarov, R.; Chang, J.-C.; Pandolfini, L.; Kawashima, T.; Tagami, M.; Nobusada, T.; Kouno, T.; Alfonso Gonzalez, C.; Albanese, R.; Dossena, F.; Haberman, N.; Ozaki, K.; Kasukawa, T.; Lenhard, B.; Frith, M.; Bodega, B.; Nicassio, F.; Calviello, L.; Bienko, M.; Legnini, I.; Hilgers, V.; Gustincich, S.; Goeke, J.; Lecellier, C. H.; Shin, J. W.; Hon, C.-C.; Carninci, P.
Show abstract
Long-read sequencing has emerged as a powerful tool for uncovering novel transcripts and genes. However, existing protocols often lack confidence in identifying the transcription start site (TSS) and fail to capture non-poly(A) RNA, thereby limiting the discovery of novel genes, particularly long non-coding RNAs (lncRNAs). In this study, we introduce Cap-trap full-length cDNA sequencing (CFC-seq), a comprehensive protocol that combines Cap-trapping and poly(A)-tailing with Oxford Nanopore sequencing. This protocol enables precise identification of TSSs and full-length transcripts. Applying CFC-seq to two in vitro differentiation time courses resulted in approximately 236 million mappable reads. The transcript Start-site Aware Long-read Assembler (SALA) was developed for de novo assembling the transcript models, leading to the identification of 39,425 confident novel genes. Using this dataset, enhancer-derived ncRNAs were re-defined with longer length and more splicing activity, which were correlated with enhancer structure. Compared to enhancers with CpG islands, TATA box enhancers were shown to be more cell type specific with fewer chromatin interaction but produced longer and more stable polyadenylated RNA. A significant proportion of these TATA box-derived eRNAs originated from LTR transposable elements. Overall, this study systematically annotated [~]24,000 novel eRNA genes and correlated their transcription properties with enhancer structure. HighlightsO_LIFrom 236 million long-reads, CFC-seq identified 39,425 novel genes with genuine TSS support. These include [~]24,000 eRNA genes. C_LIO_LISALA, a long-read assembler, was developed to facilitate genuine TSS incorporation. C_LIO_LICompared to TATA box enhancers, CGI enhancers are more ubiquitous, enriched with repressive histone mark, with more chromatin connection and are enriched in 2D and super enhancer. C_LIO_LIeRNAs derived from TATA box are longer, more stable, frequently spliced with high splicing efficiency, frequently polyadenylated, and are enriched with LTR retrotransposons. C_LIO_LIThe 3end of non-poly(A) eRNA reveal the cleavage position depleted of secondary structure. C_LI
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Functional classification of noncoding RNAs associated with distinct histone modifications by PIRCh-seq 96%
- MicroExonator enables systematic discovery and quantification of microexons across mouse embryonic development 95%
- Antisense transcription can induce expression memory via stable promoter repression 95%
Similar papers in this journal
- Assessing genome-wide dynamic changes in enhancer activity during early mESC differentiation by FAIRE-STARR-seq 96%
- Endogenous retroviruses co-opted as divergently transcribed regulatory elements shape the regulatory landscape of embryonic stem cells 96%
- Studying RNA#8211;DNA interactome by Red-C identifies noncoding RNAs associated with repressed chromatin compartment and reveals transcription dynamics 96%
Similar papers in this journal
- Alternative splicing modulation by G-quadruplexes 96%
- Cas13d-mediated isoform-specific RNA knockdown with a unified computational and experimental toolbox 95%
- Semi-quantitative detection of pseudouridine modifications and type I/II hypermodifications in human mRNAs using direct and long-read sequencing 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.