Back

Evaluating the role of pre-training dataset size and diversity on single-cell foundation model performance

DenAdel, A.; Hughes, M.; Thoutam, A.; Gupta, A.; Navia, A. W.; Fusi, N.; Raghavan, S.; Winter, P. S.; Amini, A. P.; Crawford, L.

2024-12-17 genomics
10.1101/2024.12.13.628448 bioRxiv
Show abstract

The success of transformer-based foundation models on natural language and images has motivated their use in single-cell biology. Single-cell foundation models have been trained on increasingly larger transcriptomic datasets, scaling from initial studies with 1 million cells to newer atlases with over 100 million cells. This study investigates the role of pre-training dataset size and diversity on the performance of single-cell foundation models on both zero-shot and fine-tuned tasks. Using a large corpus of 22.2 million cells, we pre-train a total of 400 models, which we evaluate by conducting 6,400 experiments. Our results show that current methods tend to plateau in performance with pre-training datasets that are only a fraction of the size of current training corpora.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.