Assessing the limits of zero-shot foundation models in single-cell biology
Kedzierska, K. Z.; Crawford, L.; Amini, A. P.; Lu, A. X.
Show abstract
The advent and success of foundation models such as GPT has sparked growing interest in their application to single-cell biology. Models like Geneformer and scGPT have emerged with the promise of serving as versatile tools for this specialized field. However, the efficacy of these models, particularly in zero-shot settings where models are not fine-tuned but used without any further training, remains an open question, especially as practical constraints require useful models to function in settings that preclude fine-tuning (e.g., discovery settings where labels are not fully known). This paper presents a rigorous evaluation of the zero-shot performance of these proposed single-cell foundation models. We assess their utility in tasks such as cell type clustering and batch effect correction, and evaluate the generality of their pretraining objectives. Our results indicate that both Geneformer and scGPT exhibit limited reliability in zero-shot settings and often underperform compared to simpler methods. These findings serve as a cautionary note for the deployment of proposed single-cell foundation models and highlight the need for more focused research to realize their potential.2
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 97%
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 96%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
Similar papers in this journal
Similar papers in this journal
- Iterative point set registration for aligning scRNA-seq data 97%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 95%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.