Back

When Task-Specific Learning Outperforms Transfer Learning: A Benchmark of Gene and Expression Encoding Strategies

Sadalski, I.

2025-12-09 systems biology
10.64898/2025.12.05.690830 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWSingle-cell foundational models have emerged as a powerful tool for learning generalizable cellular representations from large-scale data. Most models in this domain use transformer backbones, which require careful engineering of gene and expression encoding strategies, yet there is no consensus on which encoding techniques are effective. While benchmarking efforts up to date have focused on evaluating downstream applications using already pretrained models, we take a fundamentally different approach: we isolate different encoding paradigms and systematically compare them by training models from scratch under controlled conditions. Moreover, we scale pretraining to 10 million cells across 100 diverse datasets, a tenfold increase compared to similar studies. Through empirical experiments, we find that contrary to common assumptions, pretrained embeddings from large protein models like ESM-2 consistently underperformed task-specific learned embeddings. Our work provides clear empirical guidance for model design decisions and establishes a systematic benchmark for evaluating encoding strategies in single-cell foundational models.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.