Back

Scalable and universal prediction of cellular phenotypes

Ji, Y.; Tejada-Lapuerta, A.; Schmacke, N. A.; Zheng, Z.; Zhang, X.; Khan, S.; Rothenaigner, I.; Tschuck, J.; Hadian, K.; Theis, F. J.

2024-08-12 cell biology
10.1101/2024.08.12.607533 bioRxiv
Show abstract

Biological systems can be interrogated by perturbing individual components and observing the consequences across molecular, cellular, and phenotypic levels. The vast combinatorial space of possible perturbations and responses makes exhaustive experimentation infeasible. Recent advances in machine learning have shown that training on diverse datasets enables transfer learning across tasks, capturing patterns that generalize and improving performance on previously unseen problems. Inspired by this principle, we present Prophet, a transformer-based model pretrained on a vast, heterogeneous collection of perturbation experiments. This pretraining allows Prophet to predict the outcomes of untested genetic or chemical perturbations in novel cellular contexts, spanning phenotypes such as gene expression, viability, and morphology. By leveraging shared structure across apparently disconnected assays, Prophet provides a scalable framework for large-scale virtual screening and prioritization of informative experiments. Prophet consistently outperforms baseline models, including those trained on single phenotypes, showing that transfer learning between phenotypes not only is possible but improves predictive accuracy. Its capabilities extends to in vivo developmental systems, where it recapitulates known lineage biology and proposes new candidates. In a large-scale in silico screen for melanoma, Prophet identified and experimentally validated compounds with selective activity that mirrored clinically approved therapies, demonstrating its ability to transform perturbation biology into a predictive and scalable engine for therapeutic discovery.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.