A transcriptomics-native foundation model for universal cell representation and virtual cell synthesis
Jiang, X.; Xie, J.
Show abstract
Current single-cell foundation models rely on language-model architectures that ignore transcriptomic data distributions, often underperforming specialized methods. We introduce xVERSE, a transcriptomics-native foundation model coupling batch-invariant representation learning with the probabilistic generation of expression profiles. xVERSE outperforms the leading foundation and batch-effect correction methods in representation learning by 17.9% and 11.4%, respectively, successfully preserving biological heterogeneity while diminishing batch effects. Furthermore, xVERSE surpasses the second-best spatial imputation method by 34.3% and uniquely synthesizes virtual cells indistinguishable from biological data (AUROC{approx} 0.5). As a powerful data-augmentation engine, xVERSE utilizes these high-fidelity virtual cells to enable accurate clustering and marker detection in tiny datasets--resolving rare cell types with as few as four cells--while improving the generalizability of cross-modality predictions across diverse pathological states. These results establish xVERSE as a transformative framework unlocking analytical capabilities beyond conventional models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 98%
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 98%
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 98%
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 97%
- scAlign: a tool for alignment, integration and rare cell identification from scRNA-seq data 97%
- CMOT: Cross Modality Optimal Transport for multimodal inference 96%
Similar papers in this journal
Similar papers in this journal
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 98%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 97%
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.