Back

Raw-count embeddings improve single-cell foundation models

Schlede, S.; Muruganandan, T. P.; Gojjam Kantharaju, S.; Kisis, I.; Boecker, M.; Kim Alves Carpinteiro, M.; Schmitz, A.; Buchwald, L. M.; Sakthivelu, V.; Gülcüler Balta, G. S.; Anstötz, M.; Rueger, M. A.; Thomas, R. K.; Beleggia, F.

2026-07-03 bioinformatics
10.64898/2026.06.29.735389 bioRxiv
Show abstract

Single-cell transformer foundation models have grown to hundreds of millions of parameters, yet the preprocessing choices that underlie them, including gene ranking and library-size normalisation, have not been systematically benchmarked. Testing seven strategies, we find these elaborations are largely unnecessary: non-normalised, log-transformed counts give the best performance, and gene order barely matters, with even random ordering outperforming sophisticated rank-based schemes. The resulting model, Gene Intelligence, projects log1p-transformed raw counts directly onto each token embedding and jointly predicts masked tokens and counts, using no normalisation, positional encoding, or read-depth tokens. Despite this simplicity, it achieves state-of-the-art performance in the tested gene-level tasks and in doublet detection, and matches large current foundation models on cell-classification tasks while using 10- to 200-fold fewer parameters.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.