A unified framework enables accessible deployment and comprehensive benchmarking of single-cell foundation models
Hou, S.; Yang, P.; Ma, W.; Wang, J. X.; Zhou, X.
Show abstract
AbstractRecent years have seen rapid growth in single-cell foundation models (scFMs), raising expectations for transformative advances in genomic data analysis. However, their adoption has been hindered by inconsistent performance across datasets, fragmented software ecosystems, high technical barriers, and the lack of best practices established through systematic, reproducible benchmarks. Here we present a unified, extensible, and fully automated computational framework that standardizes the execution, evaluation, and extension of diverse scFMs. The framework harmonizes software environments, eliminates manual configuration, and enables large-scale, reproducible evaluation across heterogeneous datasets and training regimes. Leveraging this infrastructure, we systematically benchmark thirteen foundation models alongside classical baselines across more than fifty datasets under zero-shot, few-shot, and fine-tuning settings. We show that pretrained embeddings capture biologically meaningful structure and provide clear advantages in low-label and transfer-learning scenarios, while classical PCA approach remains competitive or even preferable in others. Together, this work lowers technical barriers, delivers best practices, and establishes a transparent and reproducible standard for community-wide evaluation, accelerating rigorous development and adoption of scFMs.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.