Back

Confounding effects of inferring gene co-expression networks from pooled data from different biological populations

Runghen, R.; Eliassi-Rad, T.; Bolnick, D. I.

2026-06-29 bioinformatics
10.64898/2026.06.23.734063 bioRxiv
Show abstract

Weighted Gene Co-expression Network Analysis (WGCNA) is routinely applied to pooled datasets from multiple biological populations, genotypes, or treatment groups, implicitly assuming a shared module structure across groups. While the distortion of pairwise correlations by pooling heterogeneous groups is well established statistically, three aspects of this problem have received little systematic attention in the context of co-expression network analysis: the extent to which pooling disrupts the discrete module-level community structure inferred by WGCNA; whether this disruption is detectable from the global topology metrics researchers routinely report; and how prevalent the pooling practice is in published multi-group WGCNA studies. Using analytical toy examples and a four-scenario simulation framework, we address all three questions. Module preservation Zsummary scores declined progressively with between-population divergence, from full preservation under identical populations (mean median Zsummary = 25.2 {+/-} 3.3, 95% interval 19.0--30.7 across 20 simulation replicates) to substantial disruption when both network structure and mean expression differed (mean median Zsummary = 11.9 {+/-} 1.0, 95% interval 10.2--13.5). This disruption was undetectable from global topology metrics: modularity and clustering coefficient remained stable across all scenarios, while edge density was sensitive but non-specific. These findings were corroborated in an empirical reanalysis of divergent lake and stream stickleback transcriptomes, where merged analysis collapsed 26 lake-specific and 59 stream-specific modules into only 19 merged modules. A survey of 100 publications found that 78.7% (95% CI 69.4--87.9%) of multi-group WGCNA studies with sufficient methodological reporting used a single merged analysis. Results were robust across network sizes of 250--1,000 genes and rewiring rates of 10--50%. We provide concrete recommendations including module preservation testing in both directions, population-specific baseline networks, and consensus WGCNA as a principled alternative.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 19%
9.8%
2
BMC Bioinformatics
457 papers in training set
Top 0.9%
8.0%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.8%
8.0%
4
PLOS Computational Biology
1863 papers in training set
Top 4%
8.0%
5
Methods in Ecology and Evolution
176 papers in training set
Top 0.4%
6.8%
6
Nature Communications
5641 papers in training set
Top 26%
5.6%
7
Scientific Reports
3612 papers in training set
Top 23%
4.4%
50% of probability mass above
8
Genome Biology
637 papers in training set
Top 3%
3.6%
9
Molecular Ecology Resources
171 papers in training set
Top 0.6%
3.5%
10
BMC Genomics
406 papers in training set
Top 2%
3.3%
11
PeerJ
308 papers in training set
Top 3%
3.1%
12
Ecology and Evolution
267 papers in training set
Top 3%
2.7%
13
Peer Community Journal
281 papers in training set
Top 2%
2.5%
14
Philosophical Transactions of the Royal Society B: Biological Sciences
72 papers in training set
Top 0.8%
1.5%
15
Evolutionary Applications
108 papers in training set
Top 0.9%
1.5%
16
PLOS Biology
486 papers in training set
Top 6%
1.5%
17
Genome Research
468 papers in training set
Top 4%
1.5%
18
Frontiers in Genetics
230 papers in training set
Top 3%
1.4%
19
Bioinformatics
1204 papers in training set
Top 8%
1.1%
20
Molecular Ecology
336 papers in training set
Top 3%
1.1%
21
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.9%
22
BMC Biology
265 papers in training set
Top 5%
0.9%
23
G3: Genes, Genomes, Genetics
252 papers in training set
Top 5%
0.6%
24
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 44%
0.6%
25
mSphere
302 papers in training set
Top 7%
0.6%
26
GigaScience
212 papers in training set
Top 5%
0.6%
27
mSystems
394 papers in training set
Top 6%
0.6%
28
Microbiome
154 papers in training set
Top 2%
0.6%
29
Communications Biology
993 papers in training set
Top 34%
0.6%
30
Journal of Computational Biology
48 papers in training set
Top 1%
0.6%