Back

SKIPHOS: non-kinase specific phosphorylation site prediction with random forests and amino acid skip-gram embeddings

Trac, T. Q.; Phan, K. H.; Nguyen, C. M.; Dang, H. T.

2019-10-05 bioinformatics
10.1101/793794 bioRxiv
Show abstract

MotivationPhosphorylation, which is catalyzed by kinase proteins, is in the top two most common and widely studied types of known essential post-translation protein modification (PTM). Phosphorylation is known to regulate most cellular processes such as protein synthesis, cell division, signal transduction, cell growth, development and aging. Various phosphorylation site prediction models have been developed, which can be broadly categorized as being kinase-specific or non-kinase specific (general). Unlike the latter, the former requires a large enough number of experimentally known phosphorylation sites annotated with a given kinase for training the model, which is not the case in reality: less than 3% of the phosphorylation sites known to date have been annotated with a responsible kinase. To date, there are a few non-kinase specific phosphorylation site prediction models proposed.\n\nResultsThis paper proposes SKIPHOS, a non-kinase specific phosphorylation site prediction model based on random forests on top of a continuous distributed representation of amino acids. Experimental results on the benchmark dataset and the independent test set demonstrate that SKIPHOS compares favorably to recent state-of-the-art related methods for three phosphorylation residues. Although being trained on phosphorylation sites in mamals, SKIPHOS can yield predictions for Y residues better than PHOSFER, a recently proposed plants-specific phosphorylation prediction model.\n\nAvailability and ImplementationSKIPHOS Web Server is freely available for non-commercial use at http://fit.uet.vnu.edu.vn/SKIPHOS or http://112.137.130.46:5000.\n\nContacthai.dang@vnu.edu.vn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.