Generative single-cell transcriptomics via large language models
Choi, H.; Shin, H.; Lee, D.; Lee, D.
Show abstract
Single-cell and spatial transcriptomics have generated vast atlases of cellular states, yet these data are almost exclusively used for analysis rather than generation. Here we introduce the LLM-based model PGL, Portraying Gene Language, a framework that reframes single-cell transcriptomes as a language generation modeling problem. PGL represents each cell as a long sequence of gene-expression tokens and uses a large language model to synthesize complete single-cell RNA-seq profiles from metadata alone, such as tissue and disease context. PGL-generated cells recapitulate dataset-specific transcriptomic structure, align with known cancer subtype biology, and mix coherently with real single-cell datasets. Notably, generated cells can be used as effective references for spatial transcriptomics, enabling accurate cell-type mapping without matched single-cell atlases. By shifting single-cell modeling from representation learning to cell generation, PGL enables virtual cohort construction, hypothesis generation, and reference-on-demand analysis, positioning generative language models as foundational tools for in silico single-cell biology.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Characterizing Spatially Continuous Variations in Tissue Microenvironment through Niche Trajectory Analysis 97%
- STHD: probabilistic cell typing of single Spots in whole Transcriptome spatial data with High Definition 97%
- High-precision cell-type mapping and annotation of single-cell spatial transcriptomics with STAMapper 97%
Similar papers in this journal
- Conserved epigenetic regulatory logic infers genes governing cell identity 97%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 96%
- Multiome Perturb-seq unlocks scalable discovery of integrated perturbation effects on the transcriptome and epigenome 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.