To join or not to join: handling biological replicates in long-read RNA sequencing data
Jetzinger, F.; Paniagua, A.; Cormack, S.; Mestre-Tomas, J.; Morante-Redolat, J. M.; Farinas, I.; Ferrandez-Peral, L.; Goetz, S.; Monzo, C.; Conesa, A.
Show abstract
Long-read RNA sequencing (lrRNA-seq) has revolutionized transcriptomics facilitating the study of alternative splicing and resulting in identification of thousands of novel transcripts. While isoform identification has received significant attention, the handling of biologically replicated lrRNA-seq datasets remains less explored. However, how multiple samples are combined in a lrRNA-seq study may strongly impact transcript identification. This study defines and evaluates two strategies for obtaining consensus transcriptomes from multi-sample lrRNA-seq data: "Join & Call", where reads from all samples are combined before transcript identification, and "Call & Join", where transcript identification is performed on individual samples before combining the resulting annotations. We applied these strategies to a highly replicated dataset of mouse brain and kidney tissues, using both PacBio and ONT technologies, across six widely used transcript reconstruction tools. Our results indicate that the optimal strategy depends on the chosen computational tool and research objective. We found that Join & Call is generally more suitable for discovering rarely occurring, novel isoforms, as pooling evidence increases confidence in calling lowly-expressed transcripts. Conversely, Call & Join is computationally more efficient and often preferable for highly replicated datasets when the investigation of rare novel transcripts is not the primary objective. Our findings provide a conceptual and practical framework for multi-sample transcriptome reconstruction, guiding best practices in the context of increasingly large-scale lrRNA-seq studies.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Contrasting and Combining Transcriptome Complexity Captured by Short and Long RNA Sequencing Reads 96%
- Illumina But With Nanopore: Sequencing Illumina libraries at high accuracy on the ONT MinION using R2C2 95%
- Differences in molecular sampling and data processing explain variation among single-cell and single-nucleus RNA-seq experiments 95%
Similar papers in this journal
- Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2 96%
- Biochemical-free enrichment or depletion of RNA classes in real-time during direct RNA sequencing with RISER 95%
- Quantitative analysis of C. elegans transcripts by Nanopore direct-cDNA sequencing reveals terminal hairpins in non trans-spliced mRNAs 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.