Back

ProCyon: A multimodal foundation model for protein phenotypes

Queen, O.; Huang, Y.; Calef, R.; Giunchiglia, V.; Chen, T.; Dasoulas, G.; Tai, L.; Ektefaie, Y.; Noori, A.; Brown, J.; Cobley, T.; Hrovatin, K.; Hartvigsen, T.; Theis, F.; Pentelute, B. L.; Khurana, V.; Kellis, M.; Zitnik, M.

2024-12-15 bioinformatics
10.1101/2024.12.10.627665 bioRxiv
Show abstract

Characterizing human proteins remains a major challenge: approximately 29% of human proteins lack experimentally validated functions and even well-annotated proteins often lack context-specific phenotypic insights. To enable universal modeling of protein phenotypes, we present PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW, a multimodal foundation model that utilizes protein sequence, structure, and natural language for generating and predicting protein phenotypes across diverse knowledge domains. PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW is trained on our novel dataset, PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW-IO_SCPLOWNSTRUCTC_SCPLOW, with 33 million protein phenotype instructions. On dozens of benchmarking tasks, PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW performs competitively against single-modal and multimodal models. Further, PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW conditionally retrieves proteins via mechanisms of action of small molecule drugs and disease contexts, and it generates candidate phenotypic descriptions for poorly characterized proteins, including those implicated in Parkinsons disease that were identified after PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOWs knowledge cutoff date. We experimentally confirm PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOWs predictions in multiple sclerosis using post-mortem brain RNA-seq, identifying novel MS genes and elucidating associated pathway mechanisms consistent with cortical pathology. PO_SCPLOWROC_SCPLOWCO_SCPLOWYONC_SCPLOW paves the way toward a general approach to generate functional insights into the human proteome.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.