Towards a systematic characterization of protein complex function: a natural language processing and machine-learning framework
Sharma, V. S.; Fossati, A.; Ciuffa, R.; Buljan, M.; Williams, E. G.; Chen, Z.; Shao, W.; Pedrioli, P. G. A.; Purcell, A. W.; Rodriguez Martinez, M.; Song, J.; Manica, M.; Aebersold, R.; Li, C.
Show abstract
It is a general assumption of molecular biology that the ensemble of expressed molecules, their activities and interactions determine biological processes, cellular states and phenotypes. Quantitative abundance of transcripts, proteins and metabolites are now routinely measured with considerable depth via an array of "OMICS" technologies, and recently a number of methods have also been introduced for the parallel analysis of the abundance, subunit composition and cell state specific changes of protein complexes. In comparison to the measurement of the molecular entities in a cell, the determination of their function remains experimentally challenging and labor-intensive. This holds particularly true for determining the function of protein complexes, which constitute the core functional assemblies of the cell. Therefore, the tremendous progress in multi-layer molecular profiling has been slow to translate into increased functional understanding of biological processes, cellular states and phenotypes. In this study we describe PCfun, a computational framework for the systematic annotation of protein complex function using Gene Ontology (GO) terms. This work is built upon the use of word embedding-- natural language text embedded into continuous vector space that preserves semantic relationships-- generated from the machine reading of 1 million open access PubMed Central articles. PCfun leverages the embedding for rapid annotation of protein complex function by integrating two approaches: (1) an unsupervised approach that obtains the nearest neighbor (NN) GO term word vectors for a protein complex query vector, and (2) a supervised approach using Random Forest (RF) models trained specifically for recovering the GO terms of protein complex queries described in the CORUM protein complex database. PCfun consolidates both approaches by performing the statistical test for the enrichment of the top NN GO terms within the child terms of the predicted GO terms by RF models. Thus, PCfun amalgamates information learned from the gold-standard protein-complex database, CORUM, with the unbiased predictions obtained directly from the word embedding, thereby enabling PCfun to identify the potential functions of putative protein complexes. The documentation and examples of the PCfun package are available at https://github.com/sharmavaruns/PCfun. We anticipate that PCfun will serve as a useful tool and novel paradigm for the large-scale characterization of protein complex function.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- FunCoup 6: advancing functional association networks across species with directed links and improved user experience 94%
- ChemPert: mapping between chemical perturbation and transcriptional response for non-cancer cells 93%
- MetaOmGraph: a workbench for interactive exploratory data analysis of large expression datasets 93%
Similar papers in this journal
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 94%
- MENDELSEEK: An algorithm that predicts Mendelian Genes and elucidates what makes them special 94%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.