Foundation Models Meet Imbalanced Single-Cell Data When Learning Cell Type Annotations
Alsabbagh, A. R.; Maillo Ruiz de Infante, A.; Gomez-Cabrero, D.; Kiani, N.; Khan, S. A.; Tegner, J. N.
Show abstract
With the emergence of single-cell foundation models, an important question arises: how do these models perform when trained on datasets having an imbalance in cell type distribution due to rare cell types or biased sampling? We benchmark three foundation models, scGPT, scBERT, and Geneformer, using skewed single-cell cell-type distribution for cell-type annotation. While all models had reduced performance when challenged with rare cell types, scGPT and scBERT, performed better than Geneformer. Notably, in contrast to scGPT and scBERT, Geneformer uses ordinal positions of the tokenized genes rather than actual raw gene expression values. To mitigate the effect of a skewed distribution, we find that random oversampling, but not random undersampling, improved the performance for all three foundation models. Finally, scGPT, using FlashAttention, has the fastest computational speed, whereas scBERT is more memory-efficient. We conclude that tokenization and data representation are essential areas of research, and new strategies are needed to mitigate the effects of imbalanced learning in single-cell foundation models. Code and data for reproducibility are available at https://github.com/SabbaghCodes/ImbalancedLearningForSingleCellFoundationModels.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 97%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 96%
- Graph Contrastive Learning as a Versatile Foundation for Advanced scRNA-seq Data Analysis 96%
Similar papers in this journal
Similar papers in this journal
- parSMURF, a High Performance Computing tool for the genome-wide detection of pathogenic variants 95%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 95%
- Profiling the baseline performance and limits of machine learning models for adaptive immune receptor repertoire classification 95%
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 95%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.