Back

Gene-specific exponent-corrected normalization for library size in bulk RNA-seq

Yin, R.; Li, D.; Zong, W.; Ketchesin, K. D.; Seney, M. L.; McClung, C. A.; Baldoni, P. L.; Tseng, G. C.

2026-07-09 bioinformatics
10.64898/2026.07.04.736167 bioRxiv
Show abstract

Correcting for library size is an essential step in bulk RNA-seq analyses, as differences in sequencing depth across samples can obscure biological signal with technical noise. While numerous normalization methods and model-based strategies have been proposed, we demonstrate here that library size-normalized counts and differential expression results obtained from such widely adopted approaches often remain strongly correlated with library size in large-scale RNA-seq experiments. Through a systematic analysis of over 100 publicly available GEO and TCGA RNA-seq datasets with raw count data, we show that library size association is observed for a substantial proportion of genes even after state-of-the-art library size correction approaches recommended by leading normalization tools. To address this issue, we propose gecco, a gene-specific exponent-corrected normalization method for RNA-seq counts that incorporates library size directly into the statistical framework via a gene-specific correction term, rather than applying a uniform adjustment factor across all genes. This formulation generalizes existing normalization approaches and yields normalized counts that are free of residual library size effects. Using both simulation studies and real large-scale RNA-seq datasets, we show that our method mitigates library size bias while preserving biological signal across a range of parameter settings. We further demonstrate that our approach leads to higher detection accuracy and more biologically meaningful pathway enrichment results in downstream differential expression and rhythmicity analyses without compromising false discovery rate control. Our method is implemented in R and is fully compatible with the widely used differential expression analysis methods DESeq2 and edgeR.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.1%
14.9%
2
Bioinformatics
1204 papers in training set
Top 2%
12.5%
3
Genome Biology
637 papers in training set
Top 1%
9.7%
4
Nucleic Acids Research
1281 papers in training set
Top 3%
6.7%
5
BMC Genomics
406 papers in training set
Top 1%
5.1%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
5.1%
50% of probability mass above
7
BMC Bioinformatics
457 papers in training set
Top 2%
5.1%
8
Nature Communications
5641 papers in training set
Top 31%
4.3%
9
Scientific Reports
3612 papers in training set
Top 27%
4.0%
10
Genome Research
468 papers in training set
Top 2%
3.2%
11
PLOS Computational Biology
1863 papers in training set
Top 12%
2.4%
12
Bioinformatics Advances
203 papers in training set
Top 2%
2.4%
13
Journal of Computational Biology
48 papers in training set
Top 0.5%
2.1%
14
RNA
189 papers in training set
Top 0.8%
2.1%
15
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.1%
1.7%
16
Cell Systems
201 papers in training set
Top 3%
1.7%
17
PLOS ONE
5266 papers in training set
Top 50%
1.7%
18
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
19
Nature Methods
385 papers in training set
Top 5%
1.1%
20
Genes
144 papers in training set
Top 3%
1.0%
21
Communications Biology
993 papers in training set
Top 31%
0.8%
22
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
23
Statistics in Medicine
40 papers in training set
Top 0.7%
0.6%
24
PeerJ
308 papers in training set
Top 13%
0.6%
25
GigaScience
212 papers in training set
Top 5%
0.6%