Back

Cross-chemistry single-nucleus RNA-seq identifies gene length and CpG-island promoters as determinants of transcriptional noise.

Czapiewski, R.; Chiang, M.; Ding, J.; Naughton, C.; Grimes, G. R.; Marenduzzo, D.; Gilbert, N.

2026-08-28 genomics
10.64898/2026.08.25.747056 bioRxiv
Show abstract

Cell-to-cell transcriptional heterogeneity, or noise, is an intrinsic property of the transcriptome with implications for development, disease progression, and aging. Bulk RNA-seq masks this variability by averaging gene expression across cells, whereas single-cell RNA sequencing (scRNA-seq) resolves it. Nevertheless, separating biological noise from technical variance remains challenging, particularly across platforms with different chemistries. We benchmarked two widely adopted technologies, Evercode WT (SPLiT-seq, Parse Biosciences) and Chromium (10x Genomics), on human lymphoblastoid nuclei. Evercode WT achieved targeted sequencing depth and nuclei number far more reliably, and its random-hexamer priming yielded more intronic reads and non-coding RNA genes; Chromium recovered more cells and detected polyadenylated transcripts and cell-line markers more sensitively. Despite these opposing biases, the platforms showed comparable gene detection and strongly correlated expression profiles. Using datasets from both platforms, we defined a noise metric detrended from mean expression and showed that per-gene estimates were reproducible across chemistries. Noise was lower in G2M than in G1 and was most strongly associated with gene length rather than exonic length. Expression of genes with CpG-island promoters was less variable than that of those without. This study establishes a platform-independent basis for quantifying transcriptional noise and a framework for selecting an appropriate scRNA-seq platform.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.