Back

Cell2Sentence: Teaching Large Language Models the Language of Biology

Levine, D.; Rizvi, S. A.; Levy, S.; Pallikkavaliyaveetil MohammedSheriff, N.; Wu, R.; Han, I.; Zhang, Z.; Fonseca, A.; Chen, X.; Ghadermarzi, S.; Karbasi, A.; Dhodapkar, R. M.; van Dijk, D.

2023-11-29 bioinformatics
10.1101/2023.09.11.557287 bioRxiv
Show abstract

We introduce Cell2Sentence (C2S), a novel method to directly adapt large language models to a biological context, specifically single-cell transcriptomics. By transforming gene expression data into "cell sentences," C2S bridges the gap between natural language processing and biology. We demonstrate cell sentences enable the fine-tuning of language models for diverse tasks in biology, including cell generation, complex cell-type annotation, and direct data-driven text generation. Our experiments reveal that GPT-2, when fine-tuned with C2S, can generate biologically valid cells based on cell type inputs, and accurately predict cell types from cell sentences. This illustrates that language models, through C2S fine-tuning, can acquire a significant understanding of single-cell biology while maintaining robust text generation capabilities. C2S offers a flexible, accessible framework to integrate natural language processing with transcriptomics, utilizing existing models and libraries for a wide range of biological applications.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.