Spatium: A Protein Language Foundation Model for Spatial Proteomics
Wang, T.; Wu, S.; Huang, L.; Liu, J.; Huang, K.; Zhou, X.
Show abstract
Spatial proteomics provides single-cell protein measurements under highly constrained and heterogeneous protein panels across datasets, resulting in limited and partially overlapping measurement spaces for cellular characterization. Existing analyses predominantly rely on statistical or task-specific modeling, while learning scalable representations of spatial protein data remain underexplored. This gap motivates the need for models that can learn stable representations of cellular identity from constrained protein measurements. Here we introduce Spatium, a protein language foundation model trained on over 51 million cells across multiple spatial proteomics platforms. Spatium learns intrinsic co-expression hierarchies that capture cell identity in a manner robust to panel composition and measurement scale. Spatium builds a generalizable representation of cell states grounded in biologically interpretable protein expression patterns. We demonstrate that Spatium learns biologically meaningful cell representations across multiple downstream tasks. Spatium recovers accurate cell identities with marker expression patterns consistent with known biology and reveals functionally distinct spatial microenvironments characterized by coherent marker enrichment signatures. It further enables reconstruction of missing protein measurements while preserving biologically meaningful expression patterns. Across these analyses, Spatium demonstrates stable and interpretable performance with lightweight task-specific adaptation, highlighting the robustness of the learned representations across diverse biological and experimental contexts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 98%
- Compressed sensing expands the multiplexity of imaging mass cytometry 96%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 96%
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 96%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 96%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 95%
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 96%
- STHD: probabilistic cell typing of single Spots in whole Transcriptome spatial data with High Definition 95%
- Smoother: A Unified and Modular Framework for Incorporating Structural Dependency in Spatial Omics Data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.