PICNIC identifies condensate-forming proteins across organisms
Hadarovich, A.; Singh, H. R.; Ghosh, S.; Rostam, N.; Hyman, A. A.; Toth-Petroczy, A.
Show abstract
Biomolecular condensates are membraneless organelles that can concentrate hundreds of different proteins to operate essential biological functions. However, accurate identification of their components remains challenging and biased towards proteins with high structural disorder content with focus on self-phase separating (driver) proteins. Here, we present a machine learning algorithm, PICNIC (Proteins Involved in CoNdensates In Cells) to classify proteins involved in biomolecular condensates regardless of their role in condensate formation. PICNIC successfully predicts condensate members by identifying amino acid patterns in the protein sequence and structure in addition to the intrinsic disorder and outperforms previous methods. We performed extensive experimental validation in cellulo and demonstrated that PICNIC accurately predicts 21 out of 24 condensate-forming proteins regardless of their structural disorder content. Even though increasing disorder content was associated with organismal complexity, we found no correlation between predicted condensate proteome content and disorder content across organisms. Overall, we applied a novel machine learning classifier to interrogate condensate components at single protein and whole-proteome levels across the tree of life (picnic.cd-code.org).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 96%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 96%
- hu.MAP 2.0: Integration of over 15,000 proteomic experiments builds a global compendium of human multiprotein assemblies 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.