Universal Cell Embeddings: A Foundation Model for Cell Biology
Rosen, Y.; Roohani, Y.; Agrawal, A.; Samotorcan, L.; Consortium, T. S.; Quake, S. R.; Leskovec, J.
10.1101/2023.11.28.568918 bioRxivShow abstract
Developing a universal representation of cells which encompasses the tremendous molecular diversity of cell types within the human body and more generally, across species, would be transformative for cell biology. Recent work using single-cell transcriptomic approaches to create molecular definitions of cell types in the form of cell atlases has provided the necessary data for such an endeavor. Here, we present the Universal Cell Embedding (UCE) foundation model. UCE was trained on a corpus of cell atlas data from human and other species in a completely self-supervised way without any data annotations. UCE offers a unified biological latent space that can represent any cell, regardless of tissue or species. This universal cell embedding captures important biological variation despite the presence of experimental noise across diverse datasets. An important aspect of UCEs universality is that any new cell from any organism can be mapped to this embedding space with no additional data labeling, model training or fine-tuning. We applied UCE to create the Integrated Mega-scale Atlas, embedding 36 million cells, with more than 1,000 uniquely named cell types, from hundreds of experiments, dozens of tissues and eight species. We uncovered new insights about the organization of cell types and tissues within this universal cell embedding space, and leveraged it to infer function of newly discovered cell types. UCEs embedding space exhibits emergent behavior, uncovering new biology that it was never explicitly trained for, such as identifying developmental lineages and embedding data from novel species not included in the training set. Overall, by enabling a universal representation for every cell state and type, UCE provides a valuable tool for analysis, annotation and hypothesis generation as the scale and diversity of single cell datasets continues to grow.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 98%
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 97%
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 97%
Similar papers in this journal
Similar papers in this journal
- Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of inflammation 97%
- Super-resolved spatial transcriptomics by deep data fusion 96%
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.