Context-dependent correlations mislead transcriptomic network inference in bulk and single-cell data
Asiaee, A.; Bombina, P.; McGee, R. L.; Reed, J.; Abrams, Z. B.; Abruzzo, L. V.; Coombes, K. R.
Show abstract
BackgroundCorrelation is the dominant input to co-expression module discovery and miRNA-target inference. Both rely on an implicit assumption: a Pearson coefficient pooled across heterogeneous samples, whether tissues, cancer types, or cell types, estimates one biologically meaningful quantity. Simpsons paradox makes this assumption fragile in principle, since between- group mean shifts can dominate or reverse within-group associations. How often this happens in real transcriptomic data has not been quantified. ResultsAcross 8,890 TCGA tumors from 31 cancer cohorts and 23,170,038 miRNA-mRNA pairs, 94.8% of pairs showed both positive and negative within-cohort correlations. Restricting to the high-variance domain of one million pairs, 13.3% of pooled correlations with |rglobal|[≥]0.2 reversed against the within-cohort majority at sign tolerance{varepsilon} = 0.05. Heterogeneity was the rule rather than the exception (median I2 = 0.86, IQR 0.80-0.90), and 99.5% of pairs rejected equal correlation across cohorts at FDR < 0.05. Of 692,770 experimentally validated miRTarBase v10 targets measurable in our data, only 0.9% were uniformly negative across cohorts. The pattern recurred across modalities. In GTEx, 21.0% of pooled signs disagreed with the tissue majority, and 23.5% of pairs flipped sign after tissue-mean removal. In 10x PBMC scRNA-seq, 13.1% of gene-gene correlations flipped after cell-type-mean removal; in CITE-seq, 37.9% of protein-RNA pairs flipped under a joint WNN partition of cells. Refining context reduced reversal, though by how much depended on the partition: within BRCA, 5.5% of pairs reversed under molecular PAM50 subtypes versus 0.35% under clinical IHC receptor status, and refining T cells into transcriptome-defined subtypes cut PBMC reversal from 11.8% to 0.13%. ConclusionsA single pooled correlation coefficient can invert direction relative to its within-context constituents at rates that are not negligible. Correlations should be reported with their context: the within-context distribution, a heterogeneity statistic, and a diagnostic that separates between-context mean shifts from within-context association. We provide a small R interface that computes these summaries.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 96%
- Cross-dataset pan-cancer detection: Correlating cell-free DNA fragment coverage with open chromatin sites across cell types 94%
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 95%
- Differential Allele-Specific Expression Uncovers Breast Cancer Genes Dysregulated By Cis Noncoding Mutations 95%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 93%
Similar papers in this journal
- Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of inflammation 95%
- Robust decomposition of cell type mixtures in spatial transcriptomics 94%
- scPrisma: inference, filtering and enhancement of periodic signals in single-cell data using spectral template matching 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.