Back

Scalable multi-group nonnegative spatial factorization for spatial genomics data with cell-type heterogeneity

Chumpitaz-Diaz, L.; Shrestha, P.; Engelhardt, B. E.

2026-07-03 genomics
10.64898/2026.06.29.735224 bioRxiv
Show abstract

Spatial transcriptomics (ST) technologies enable the study of gene expression within the spatial context of tissues, providing insights into tissue structure, cellular interactions, and disease progression. However, existing dimension reduction methods often overlook spatial information or struggle to distinguish spatial gene patterns from those driven by cell-type differences, limiting biological interpretability by convolving differences in gene expression patterns with differences in cell-type proportions. To address these challenges, we introduce the scalable multi-group nonnegative spatial factorization (smNSF), a computationally-tractable probabilistic framework that integrates spatial coordinates and cell-type labels into a unified matrix factorization model. By using multi-group Gaussian processes (MGGPs) as priors, our model captures complex spatial variation in a cell-type specific way while enforcing nonnegativity to enhance interpretability. We develop a variational inference framework for MGGPs that supports scalable optimization and improves the numerical stability of smNSF. Across seven spatial transcriptomics datasets spanning diverse technologies and tissues, smNSF recovers sparse, interpretable spatial factors and, through its cell-type conditional posteriors, organizes them into cell-type enriched, cell-type specific, and universal spatial programs that are not apparent from marginal factors alone. Given cell-type labels in ST data, smNSF enables cell-type aware spatial decompositions and supports cell-type conditional posteriors for in silico exploration of relationships between spatial patterns and cellular identity.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 9%
18.4%
2
Genome Biology
637 papers in training set
Top 0.3%
18.4%
3
PLOS Computational Biology
1863 papers in training set
Top 5%
7.2%
4
Cell Systems
201 papers in training set
Top 1%
4.3%
5
Nature Methods
385 papers in training set
Top 2%
4.0%
50% of probability mass above
6
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 17%
3.2%
7
Biometrics
23 papers in training set
Top 0.1%
3.2%
8
Nature Computational Science
55 papers in training set
Top 0.3%
2.4%
9
GENETICS
483 papers in training set
Top 2%
2.4%
10
Nucleic Acids Research
1281 papers in training set
Top 8%
2.1%
11
Genome Research
468 papers in training set
Top 4%
1.7%
12
The Annals of Applied Statistics
19 papers in training set
Top 0.1%
1.7%
13
Cell
431 papers in training set
Top 6%
1.7%
14
Nature Biotechnology
172 papers in training set
Top 2%
1.7%
15
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.7%
16
Bioinformatics
1204 papers in training set
Top 7%
1.5%
17
Cell Reports Methods
165 papers in training set
Top 2%
1.5%
18
eLife
5828 papers in training set
Top 53%
1.4%
19
Biostatistics
24 papers in training set
Top 0.2%
1.3%
20
Nature Genetics
286 papers in training set
Top 4%
1.3%
21
Nature Machine Intelligence
70 papers in training set
Top 2%
1.3%
22
Cell Reports
1498 papers in training set
Top 23%
1.1%
23
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
24
Communications Biology
993 papers in training set
Top 28%
0.9%
25
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.2%
0.8%
26
PNAS Nexus
159 papers in training set
Top 3%
0.8%
27
PLOS ONE
5266 papers in training set
Top 62%
0.8%
28
PLOS Genetics
862 papers in training set
Top 12%
0.8%
29
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
30
Nature Neuroscience
252 papers in training set
Top 5%
0.8%