CELLama: Foundation Model for Single Cell and Spatial Transcriptomics by Cell Embedding Leveraging Language Model Abilities
Choi, H.; Park, J.; Kim, S.; Kim, J.; Lee, D.; Bae, S.; Shin, H.; Lee, D.
Show abstract
Large-scale single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) have transformed biomedical research into a data-driven field, enabling the creation of comprehensive data atlases. These methodologies facilitate detailed understanding of biology and pathophysiology, aiding in the discovery of new therapeutic targets. However, the complexity and sheer volume of data from these technologies present analytical challenges, particularly in robust cell typing, integration and understanding complex spatial relationships of cells. To address these challenges, we developed CELLama (Cell Embedding Leverage Language Model Abilities), a framework that leverage language model to transform cell data into sentences that encapsulate gene expressions and metadata, enabling universal cellular data embedding for various analysis. CELLama, serving as a foundation model, supports flexible applications ranging from cell typing to the analysis of spatial contexts, independently of manual reference data selection or intricate dataset-specific analytical workflows. Our results demonstrate that CELLama has significant potential to transform cellular analysis in various contexts, from determining cell types across multi-tissue atlases and their interactions to unraveling intricate tissue dynamics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 97%
- A universal approach for integrating super large-scale single-cell transcriptomes by exploring gene rankings 96%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 96%
Similar papers in this journal
- SPOTlight:Seeded NMF regression to Deconvolute Spatial Transcriptomics Spots with Single-Cell Transcriptomes 97%
- Disentangling single-cell omics representation with a power spectral density-based feature extraction 96%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.