Representing cells as sentences enables natural-language processing for single-cell transcriptomics
Dhodapkar, R. M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWGene expression matrices commonly used in single-cell transcriptomics, cannot be directly analyzed with tools developed for natural languages. By restructuring these matrices as abundance-ordered sequences of genes, we generate cell sentences: rank-normalized, positionally encoded expression data. We show that these cell sentences can be analyzed using existing tools from natural language processing to unify cell and gene representations across species.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- The axes of biology: a novel axes-based network embedding paradigm to decipher the functional mechanisms of the cell. 94%
- Discovering paracrine regulators of cell type composition from spatial transcriptomics using SPER 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.