Accurate quantification of single-nucleus and single-cell RNA-seq transcripts
Eldjarn Hjörleifsson, K.; Sullivan, D. K.; Holley, G.; Melsted, P.; Pachter, L.
Show abstract
In single-cell and single-nucleus RNA sequencing, the coexistence of nascent (unprocessed) and mature (processed) mRNA poses challenges in accurate read mapping and the interpretation of count matrices. The traditional transcriptome reference, defining the region of interest in bulk RNA-seq, restricts its focus to mature mRNA transcripts. This restriction leads to two problems: reads originating outside of the region of interest are prone to mismapping within this region, and additionally, such external reads cannot be matched to specific transcript targets. Expanding the region of interest to encompass both nascent and mature mRNA transcript targets provides a more comprehensive framework for RNA-seq analysis. Here, we introduce the concept of distinguishing flanking k-mers (DFKs) to improve mapping of sequencing reads. We have developed an algorithm to identify DFKs, which serve as a sophisticated background filter, enhancing the accuracy of mRNA quantification. This dual strategy of an expanded region of interest coupled with the use of DFKs enhances the precision in quantifying both mature and nascent mRNA molecules, as well as in delineating reads of ambiguous status.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transcriptome assembly from long-read RNA-seq alignments with StringTie2 98%
- Enhancing transcriptome expression quantification through accurate assignment of long RNA sequencing reads with TranSigner 97%
- Identifying and quantifying isoforms from accurate full-length transcriptome sequencing reads with Mandalorion 95%
Similar papers in this journal
- tinyRNA: precision analysis of small RNA-seq data with user-defined hierarchical selection rules 94%
- K2R: Tinted de Bruijn Graphs implementation for efficient read extraction from sequencing datasets 94%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.