Back

Hidden assumptions in nascent RNA sequencing pipelines define reproducibility states

Zhou, X.; Feng, C.; Zhao, Y.

2026-07-17 bioinformatics
10.64898/2026.07.13.738089 bioRxiv
Show abstract

Reproducibility of sequencing analyses is often assumed when identical data are processed with established pipelines, yet outcomes can depend on library assumptions that are not explicit to users. Here we examined commonly used pipelines for nascent RNA sequencing. Across public human PRO-seq datasets, identical inputs generated structured divergence in transcriptional profiles. Diagnostic processing combinations traced this divergence to interactions between paired-end library design, UMI organization and pipeline-embedded assumptions for read trimming, alignment and signal generation. This pattern persisted in independent human and pig PRO-seq libraries sharing a dual-end UMI design, reflecting pipeline-defined assumptions not fully accessible through user-specified parameters. Beyond PRO-seq, GRO-seq analyses showed that assay-specific library architecture can distort positional signal profiles without UMI processing, whereas PRO-cap and reannotated PRO-seq datasets showed that incomplete metadata can prevent pipeline execution or cause silent signal loss. Together, these results define reproducibility states shaped by library design, pipeline assumptions and metadata availability.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.