Back

Stack: In-Context Learning of Single-Cell Biology

Dong, M.; Adduri, A.; Gautam, D.; Carpenter, C.; Shah, R.; Ricci-Tam, C.; Kluger, Y.; Burke, D. P.; Roohani, Y. H.

2026-01-09 bioinformatics
10.64898/2026.01.09.698608 bioRxiv
Show abstract

Single-cell transcriptomics offers the promise of measuring the diversity of cellular phenotypes across species, diseases, and other biological conditions. Recently, foundation models have emerged to identify this variation, yet most methods represent each cell independently, despite technical limitations that reduce measurement precision at the single-cell level. Here, we present SO_SCPLOWTACKC_SCPLOW, a foundation model trained on 149 million uniformly preprocessed human single cells that leverages tabular attention to generate representations for each cell informed by the cells in its context. SO_SCPLOWTACKC_SCPLOW offers substantial improvements for downstream tasks in the zero-shot setting compared to baselines, whether they are zero-shot, fine-tuned, or trained from scratch on the target dataset. SO_SCPLOWTACKC_SCPLOW can perform in-context learning from unlabeled cells representing arbitrary conditions, such as a chemical perturbation or a different donor, and predict the effect of those conditions on a target cell population without requiring data-specific fine-tuning. We apply SO_SCPLOWTACKC_SCPLOW to generate Perturb Sapiens, the first human whole-organism atlas of perturbed cells, spanning 28 tissues, 40 cell classes, and 201 perturbations. We validated subsets of Perturb Sapiens using in vitro stimulation profiles. Overall, SO_SCPLOWTACKC_SCPLOW presents a new modeling framework where cells themselves act as guiding examples at inference time, unlocking general-purpose in-context learning capabilities for single-cell biology.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.