Scaling Large Language Models for Next-Generation Single-Cell Analysis
Rizvi, S. A.; Levine, D.; Patel, A.; Zhang, S.; Wang, E.; He, S.; Zhang, D.; Tang, C.; Lyu, Z.; Darji, R.; Li, M.; Sun, E.; Jeong, D.; Zhao, L.; Kwan, J.; Braun, D.; Hafler, B.; Ishizuka, J.; Dhodapkar, R.; Chung, H.; Azizi, S.; Perozzi, B.; van Dijk, D.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWSingle-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual "cell sentences," to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. Scaling the model to 27 billion parameters yields consistent improvements in predictive and generative capabilities and supports advanced downstream tasks that require synthesis of information across multi-cellular contexts. Targeted fine-tuning with modern reinforcement learning techniques produces strong performance in perturbation response prediction, natural language interpretation, and complex biological reasoning. This predictive strength enabled a dual-context virtual screen that nominated the kinase inhibitor silmitasertib (CX-4945) as a candidate for context-selective upregulation of antigen presentation. Experimental assessment in human cell models unseen during training supported this prediction, demonstrating that C2S-Scale can effectively guide the discovery of context-conditioned biology. C2S-Scale unifies transcriptomic and textual data at unprecedented scales, surpassing both specialized single-cell models and general-purpose LLMs to provide a platform for next-generation single-cell analysis and the development of "virtual cells."
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 97%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 96%
- scTrace+: enhance the cell fate inference by integrating the lineage-tracing and multi-faceted transcriptomic similarity information 95%
Similar papers in this journal
- scAlign: a tool for alignment, integration and rare cell identification from scRNA-seq data 96%
- CMOT: Cross Modality Optimal Transport for multimodal inference 96%
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 96%
Similar papers in this journal
- Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of inflammation 96%
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 96%
- scvi-tools: a library for deep probabilistic analysis of single-cell omics data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.