Back

A perturbation proteomics-based foundation model for virtual cell construction

Sun, R.; qian, l.; Li, Y.; Cheng, H.; Xue, Z.; zhang, x.; Tan, L.; Zhan, Y.; hu, w.; Xiao, Q.; Liu, Z.; zhang, g.; E, W.; Zhou, P.; Wen, H.; Zhu, Y. J.; Guo, T.

2025-02-10 systems biology
10.1101/2025.02.07.637070 bioRxiv
Show abstract

Building a virtual cell requires comprehensive understanding of protein network dynamics of a cell which necessitates large-scale perturbation proteome data and intelligent computational models learned from the proteome data corpus. Here, we generate a large-scale dataset of over 38 million perturbed protein measurements in breast cancer cell lines and develop a neural ordinary differential equation-based foundation model, namely ProteinTalks. During pretraining, ProteinTalks gains a fundamental understanding of cellular protein network dynamics. Our model encodes protein networks and exhibits consistently improved predictive accuracy across various downstream tasks, highlighting its generalization capabilities and adaptability. In cancer cells, ProteinTalks robustly predicts drug efficacy and synergy, identifies novel drug combinations, and, through its interpretability, uncovers resistance-associated proteins. When applied to more complex system, patient-derived tumor xenografts, ProteinTalks predicts potential responses to drugs. Its integration with clinical patient data enhances the prognosis prediction of breast cancer patients. Collectively, we present a foundational model based on proteome dynamics, offering potential for various downstream applications, including drug discovery, and providing a basis for developing virtual cells.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.