Back

HEIMDALL: A Modular Framework for Tokenization in Single-Cell Foundation Models

Haber, E.; Alam, S.; Ho, N.; Liu, R.; Trop, E.; Liang, S.; Yang, M.; Krieger, S.; Ma, J.

2025-11-10 bioinformatics
10.1101/2025.11.09.687403 bioRxiv
Show abstract

Foundation models trained on single-cell RNA-sequencing (scRNA-seq) data have rapidly become powerful tools for single-cell analysis. Their performance, however, depends critically on how cells are tokenized into model inputs - a design space that remains poorly understood. Here, we present HO_SCPLOWEIMDALLC_SCPLOW, a comprehensive framework and open-source toolkit for systematically evaluating tokenization strategies in single-cell foundation models (scFMs). HO_SCPLOWEIMDALLC_SCPLOW decomposes each scFM into modular components: a gene identity encoder (FG), an expression encoder (FE), and a "cell sentence" constructor (FC) with submodules (ORDER, SEQUENCE, and REDUCE) enabling fine-grained control and attribution. Using a transformer trained from scratch, we evaluate tokenization strategies for cell type classification across challenging transfer learning settings - cross-tissue, cross-species, and spatial gene-panel shifts - and separately assess reverse perturbation prediction. Tokenization choices show minimal impact in-distribution but are decisive under distribution shift, with FG and ORDER driving the largest gains and FE providing additional improvements. HO_SCPLOWEIMDALLC_SCPLOW further shows how existing strategies can be recombined to enhance generalization. By standardizing evaluation and providing an extensive library, HO_SCPLOWEIMDALLC_SCPLOW establishes a foundation for reproducible, systematic exploration of single-cell tokenization and accelerates the development of next-generation scFMs.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.