Back

scYeast: a Biological-knowledge-guided Foundation Model on Yeast Single-Cell Transcriptomics

Fan, X.; Liao, W.; Xiao, L.; Yan, X.; Lu, H.

2025-10-27 bioinformatics
10.1101/2025.08.20.671179 bioRxiv
Show abstract

Though large-scale pre-trained models are vital for foundational cell modeling, most focus on human or mouse systems, neglecting model organisms like yeast (Saccharomyces cerevisiae) and often failing to use existing biological prior knowledge effectively. To address this, we present scYeast, the first foundational cell model tailored for yeast that seamlessly embeds biological priors. scYeast employs a novel asymmetric parallel architecture to infuse transcriptional regulatory information directly into the Transformers attention mechanism, leveraging established biological knowledge during training. Pre-trained on large-scale yeast single-cell transcriptomics data, scYeast demonstrates strong generalization and biological interpretability. It excels in zero-shot tasks, such as inferring regulatory relationships and identifying critical cell states. After fine-tuning, scYeast performs exceptionally across a diverse set of tasks from cell type classification to predicting growth doubling time and gene perturbation response. Additionally, using transfer learning, scYeast can be adapted to other omics datasets, such as proteomics, thus broadening its utility in big data analysis. Overall, scYeast is a powerful tool for yeast single-cell biology research and sets a new standard for integrating foundational models with biological prior knowledge, dramatically accelerating the pace of discovery in yeast synthetic and systems biology and providing a replicable framework for other organisms.

Published in Synthetic and Systems Biotechnology · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.