Back

OmniCell: Unified Foundation Modeling of Single-Cell and Spatial Transcriptomics for Cellular and Molecular Insights

Pang, J.; Qiu, P.; He, Y.; Li, B.; Deng, Y.; Wang, J.; Lin, A.; Cao, L.; Teng, F.; Wang, H.; Fang, S.; Li, S.; Deng, Z.; Zhang, Y.; Li, Y.; li, s.; Xu, X.

2025-12-29 bioinformatics
10.64898/2025.12.29.696804 bioRxiv
Show abstract

Single-cell RNA sequencing (scRNA-seq) enables characterization of cellular heterogeneity but lacks spatial context, while Spatially Transcriptomics maps gene expression in tissues with limited single-cell resolution. Integrating the complementary strengths of these data into a unified framework remains challenging. Here, we present OmniCell, a foundation model for single-cell and spatial transcriptomics, pretrained on a large-scale corpus of 67 million single-cell and spatial transcriptomic profiles, enabling the unified multi-omics representation learning. As the first foundation model to jointly capture intra-cellular gene expression relationships and inter-cellular spatial dependencies within a unified framework, OmniCell explicitly represents tissue spatial topology by serializing spatially adjacent cells during input construction. Leveraging this unified modeling paradigm, OmniCell generates unified representations of genes, cells, and tissue spatial organization. In zero-shot evaluations, it reliably recovers cell-type structure and gene expression patterns, reconstructs co-expression relationships, and outperforms existing methods across all evaluated tasks, including cell-type deconvolution and spatial domain delineation. Applied to real spatial datasets, OmniCell resolves transitional zones at tumor margins and reveals associated inflammatory activation and immune-cell enrichment, demonstrating its capacity for high-resolution spatial profiling.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.