Unsupervised analysis of multi-experiment transcriptomic patterns with SegRNA identifies unannotated transcripts
Mendez, M.; FANTOM Consortium Main Contributors, ; Scott, M. S.; Hoffman, M. M.
Show abstract
BackgroundExploratory analysis of complex transcriptomic data presents multiple challenges. Many methods often rely on preexisting gene annotations, impeding identification and characterization of new transcripts. Even for a single cell type, comprehending the diversity of RNA species transcribed at each genomic region requires combining multiple datasets, each enriched for specific types of RNA. Currently, examining combinatorial patterns in these data requires time-consuming visual inspection using a genome browser. MethodWe developed a new segmentation and genome annotation (SAGA) method, SegRNA, that integrates data from multiple transcriptome profiling assays. SegRNA identifies recurring combinations of signals across multiple datasets measuring the abundance of transcribed RNAs. Using complementary techniques, SegRNA builds on the Segway SAGA framework by learning parameters from both the forward and reverse DNA strands. SegRNAs unsupervised approach allows exploring patterns in these data without relying on pre-existing transcript models. ResultsWe used SegRNA to generate the first unsupervised transcriptome annotation of the K562 chronic myeloid leukemia cell line, integrating multiple types of RNA data. Combining RNA-seq, CAGE, and PRO-seq experiments together captured a diverse population of RNAs throughout the genome. As expected, SegRNA annotated patterns associated with gene components such as promoters, exons, and introns. Additionally, we identified a pattern enriched for novel small RNAs transcribed within intergenic, intronic, and exonic regions. We applied SegRNA to FANTOM6 CAGE data characterizing 285 lncRNA knockdowns. Overall, SegRNA efficiently summarizes diverse multi-experiment data.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Quality control and processing of nascent RNA profiling data 96%
- ZetaSuite, A Computational Method for Analyzing Multi-dimensional High-throughput Data, Reveals Genes with Opposite Roles in Cancer Dependency 95%
- RATTLE: Reference-free reconstruction and quantification of transcriptomes from Nanopore sequencing 95%
Similar papers in this journal
- Biochemical-free enrichment or depletion of RNA classes in real-time during direct RNA sequencing with RISER 96%
- Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2 96%
- Semi-quantitative detection of pseudouridine modifications and type I/II hypermodifications in human mRNAs using direct and long-read sequencing 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.