Evaluating the role of pre-training dataset size and diversity on single-cell foundation model performance
DenAdel, A.; Hughes, M.; Thoutam, A.; Gupta, A.; Navia, A. W.; Fusi, N.; Raghavan, S.; Winter, P. S.; Amini, A. P.; Crawford, L.
Show abstract
The success of transformer-based foundation models on natural language and images has motivated their use in single-cell biology. Single-cell foundation models have been trained on increasingly larger transcriptomic datasets, scaling from initial studies with 1 million cells to newer atlases with over 100 million cells. This study investigates the role of pre-training dataset size and diversity on the performance of single-cell foundation models on both zero-shot and fine-tuned tasks. Using a large corpus of 22.2 million cells, we pre-train a total of 400 models, which we evaluate by conducting 6,400 experiments. Our results show that current methods tend to plateau in performance with pre-training datasets that are only a fraction of the size of current training corpora.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancement of network architecture alignment in comparative single-cell studies 98%
- Sampling from Disentangled Representations of Single-Cell Data Using Generative Adversarial Networks 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
Similar papers in this journal
- Simultaneous dimensionality reduction and integration for single-cell ATAC-seq data using deep learning 97%
- Delineating the Effective Use of Self-Supervised Learning in Single-Cell Genomics 96%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 96%
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 98%
- multiDGD: A versatile deep generative model for multi-omics data 97%
- scSemiProfiler: Advancing Large-scale Single-cell Studiesthrough Semi-profiling with Deep Generative Models andActive Learning 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.