Back

DeltaNMF: A Two-Stage Neural NMF for Differential Gene Program Discovery

Karpurapu, A.; Gersbach, C. A.; Singh, R.

2026-01-24 bioinformatics
10.64898/2026.01.22.701049 bioRxiv
Show abstract

Non-negative matrix factorization (NMF) is a foundational dimensionality-reduction method in single-cell transcriptomics, valued for its interpretable gene programs. However, in case-control settings common in perturbation and disease research, standard NMF conflates quantitative shifts in program usage with qualitative emergence of novel programs, obscuring whether observed differences reflect altered activity of shared programs or genuinely new cellular states. Here, we present DeltaNMF, a neural network-based reformulation of NMF that enables flexible biological priors and structured multi-stage fitting. We first demonstrate how our framework incorporates foundation model-derived gene similarities through graph-Laplacian regularization, improving program coherence by 2.4-fold on protein interaction networks. Building on this flexibility, we introduce a two-stage architecture that explicitly separates baseline programs (learned from control cells) from case-specific programs, disambiguating altered usage from novel program emergence. Our GPU-accelerated implementation achieves 20x speedup over consensus NMF while maintaining comparable accuracy. On synthetic data with known ground truth, DeltaNMF correctly identifies which programs are new versus shifted. Applied to coronary artery disease (CAD) Perturb-seq data, DeltaNMF isolates disease programs including the cerebral cavernous malformations pathway identified by Schnitzler et al., while clearly distinguishing them from baseline endothelial programs with altered activity. Our neural network formulation of NMF opens new directions for incorporating diverse biological priors into interpretable single-cell analysis, providing a principled framework for differential program discovery in case-control studies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.