Hidden assumptions in nascent RNA sequencing pipelines define reproducibility states
Zhou, X.; Feng, C.; Zhao, Y.
Show abstract
Reproducibility of sequencing analyses is often assumed when identical data are processed with established pipelines, yet outcomes can depend on library assumptions that are not explicit to users. Here we examined commonly used pipelines for nascent RNA sequencing. Across public human PRO-seq datasets, identical inputs generated structured divergence in transcriptional profiles. Diagnostic processing combinations traced this divergence to interactions between paired-end library design, UMI organization and pipeline-embedded assumptions for read trimming, alignment and signal generation. This pattern persisted in independent human and pig PRO-seq libraries sharing a dual-end UMI design, reflecting pipeline-defined assumptions not fully accessible through user-specified parameters. Beyond PRO-seq, GRO-seq analyses showed that assay-specific library architecture can distort positional signal profiles without UMI processing, whereas PRO-cap and reannotated PRO-seq datasets showed that incomplete metadata can prevent pipeline execution or cause silent signal loss. Together, these results define reproducibility states shaped by library design, pipeline assumptions and metadata availability.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comprehensive Benchmarking of CITE-seq versus DOGMA-seq Single Cell Multimodal Omics 95%
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 95%
- Quality control and processing of nascent RNA profiling data 95%
Similar papers in this journal
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 96%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 96%
- Multiplexed transcriptome discovery of RNA binding protein binding sites by antibody-barcode eCLIP 95%
Similar papers in this journal
- Multiome Perturb-seq unlocks scalable discovery of integrated perturbation effects on the transcriptome and epigenome 95%
- Iterative deep learning-design of human enhancers exploits condensed sequence grammar to achieve cell type-specificity 95%
- Direct analysis of ribosome targeting illuminates thousand-fold regulation of translation initiation 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.