When Task-Specific Learning Outperforms Transfer Learning: A Benchmark of Gene and Expression Encoding Strategies
Sadalski, I.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWSingle-cell foundational models have emerged as a powerful tool for learning generalizable cellular representations from large-scale data. Most models in this domain use transformer backbones, which require careful engineering of gene and expression encoding strategies, yet there is no consensus on which encoding techniques are effective. While benchmarking efforts up to date have focused on evaluating downstream applications using already pretrained models, we take a fundamentally different approach: we isolate different encoding paradigms and systematically compare them by training models from scratch under controlled conditions. Moreover, we scale pretraining to 10 million cells across 100 diverse datasets, a tenfold increase compared to similar studies. Through empirical experiments, we find that contrary to common assumptions, pretrained embeddings from large protein models like ESM-2 consistently underperformed task-specific learned embeddings. Our work provides clear empirical guidance for model design decisions and establishes a systematic benchmark for evaluating encoding strategies in single-cell foundational models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.