Back

scROMA: batch-aware pathway-activity inference and a ground-truth simulation framework for single-cell transcriptomics

Zhubanchaliyev, A.; Najm, M.; Laigle, V.; Bonnet, E.; Martignetti, L.

2026-08-10 systems biology
10.64898/2026.08.07.743516 bioRxiv
Show abstract

BackgroundPathway-activity analysis summarizes gene-level single-cell measurements into interpretable functional modules, but widely used methods lack an integrated significance framework, do not account for the batch effects that pervade multi-sample studies, and are not natively interoperable with Python-based workflows. The field also lacks simulation resources with ground-truth pathway activity for quantitative benchmarking. ResultsWe present scROMA, a singular-value-decomposition-based method that quantifies pathway activity as coordinated variation, with per-cell scores, per-gene contributions, and permutation-based significance, natively integrated with the Scanpy/AnnData ecosystem. Its batch-aware extension is, to our knowledge, the first to correct batch effects within the gene-set subspace rather than across the full transcriptome, isolating technical variation at the pathway level while preserving signal in other genes. We also release a generative simulation framework producing synthetic data with fully specified ground-truth activities. On simulated benchmarks scROMA is competitive across tasks, and under batch effects its batch-aware mode recovers ordinal pathway structure that full-transcriptome integration misses. Across cystic fibrosis airway, intestinal-organoid, breast cancer, and lung cancer datasets it recovers established biology while separating it from technical and inter-donor variation; in the intestinal-organoid atlas it reproducibly recovers an inflammatory program across donors, separates its sustained from transient components, and resolves cell-type-specific niche-factor targets. ConclusionsscROMA is open-source and released with the simulation framework and pre-generated benchmark datasets as a community resource, providing a scalable, statistically grounded, and batch-aware approach to pathway-level analysis in single-cell transcriptomics.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.