Back

Reinforcement learning enables single-cell foundation models to learn cellular differentiation

Chang, K.-L.; Chen, H.; Liu, Z.

2025-12-12 bioinformatics
10.64898/2025.12.09.693267 bioRxiv
Show abstract

While single-cell foundation models excel at static representation learning and single-step perturbation prediction, their capacity to model and control dynamic, sequential cell state transitions remains underexplored. Here, we introduce Differentiation with Reinforcement Learning (DiRL), a framework that transforms foundation models into predictive environments for optimizing multi-step perturbation strategies. By formulating differentiation as a goal-conditioned sequential decision-making problem, DiRL trains agents to navigate the high-dimensional latent landscape of gene expression, learning policies that direct stem cells toward specific terminal fates. Evaluated on iPSC-derived organoid datasets, DiRL outperforms random perturbation baselines with a 69% win rate, with performance scaling systematically as the planning horizon increases. Beyond optimization, DiRL offers interpretable insights into the mechanics of cell fate transitions: the learned policies successfully recover known sequential regulators in Wnt and Hedgehog signaling pathways, while the value function from the critic model recapitulates biological pseudotime orderings comparable to established trajectory inference methods. These results demonstrate that coupling reinforcement learning with foundation models enables a paradigm shift from static embedding analysis to dynamic trajectory modeling, providing a powerful engine for discovering sequential perturbation strategies in regenerative medicine.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.