Back

Species-specific small models for cell type classification approach the performance of large single cell foundation models

Mahmoudabadi, G.; Krishnan, L.; Ganapathi, T.; Pearce, J.; Quake, S.; Karaletsos, T.

2026-03-18 cell biology
10.64898/2026.03.16.711196 bioRxiv
Show abstract

Accurate cross-species cell type classification remains a key evaluation task in single-cell transcriptomics. Recent foundation models trained on millions of single-cell profiles demonstrate great in-distribution and out-of-distribution performance on this task, but their large parameter counts and substantial computational costs limit accessibility and interpretability. Here, we introduce CytoType, a simple and interpretable model for cell type classification that leverages pre-trained ESM-2 protein embeddings of protein-coding transcripts. By learning linear, cell-type-specific weights over transcript embeddings, without relying on gene count information, CytoType achieves F1 scores comparable to or exceeding those of large-scale transformer-based models. We further developed ESM-Cell Embedding (ESM-CE), an even simpler variant that only averages ESM-2 embeddings across expressed genes, which also performs competitively against foundation models. Both CytoType and ESM-CE are trained on species-specific data, maintaining high accuracy when classifying cell types with orders of magnitude fewer parameters compared to larger foundation models. For example, for human tissues, the average performance gap between CytoType and the best foundation model was 0.053 F1 points while CytoType uses 10,000x fewer trainable parameters. Additionally, we quantified the contribution of ESM-2 embeddings to cell type classification tasks and demonstrated a three fold reduction in the performance gap between CytoType and the best foundation model for nine species. Finally, we show that CytoTypes learned gene weights are biologically interpretable.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.