CellPLM: Pre-training of Cell Language Model Beyond Single Cells
Wen, H.; Tang, W.; Dai, X.; Ding, J.; Jin, W.; Xie, Y.; Tang, J.
Show abstract
The current state-of-the-art single-cell pre-trained models are greatly inspired by the success of large language models. They trained transformers by treating genes as tokens and cells as sentences. However, three fundamental differences between single-cell data and natural language data are overlooked: (1) scRNA-seq data are presented as bag-of-genes instead of sequences of RNAs; (2) Cell-cell relations are more intricate and important than inter-sentence relations; and (3) The quantity of single-cell data is considerably inferior to text data, and they are very noisy. In light of these characteristics, we propose a new pre-trained model CellPLM, which takes cells as tokens and tissues as sentences. In addition, we leverage spatially-resolved transcriptomic data in pre-training to facilitate learning cell-cell relationships and introduce a Gaussian mixture prior distribution as an additional inductive bias to overcome data limitation. CellPLM is the first single-cell pre-trained transformer that encodes cell-cell relations and it consistently outperforms existing pre-trained and non-pre-trained models in diverse downstream tasks, with 100x times higher inference speed compared to existing pre-trained models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scNODE: Generative Model for Temporal Single Cell Transcriptomic Data Prediction 97%
- SAILER: Scalable and Accurate Invariant Representation Learning for Single-Cell ATAC-Seq Processing and Integration 96%
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 96%
Similar papers in this journal
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 96%
- Deep learning of gene interactions from single cell time-course expression data 95%
- Sincast: a computational framework to predict cell identities in single cell transcriptomes using bulk atlases as references 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.