Back

SeqPIP-2020: Sequence based Protein InteractionPrediction Contest

Halder, A. K.; Mollah, A. F.; Chatterjee, P.; Kole, D. K.; Basu, S.; Plewczynski, D.

2020-11-13 bioinformatics
10.1101/2020.11.12.380774 bioRxiv
Show abstract

Computational protein-protein interaction (PPI) prediction techniques can contribute greatly in reducing time, cost and false-positive interactions compared to experimental approaches. Sequence is one of the key and primary information of proteins that plays a crucial role in PPI prediction. Several machine learning approaches have been applied to exploit the characteristics of PPI datasets. However, these datasets greatly influence the performance of predicting models. So, care should be taken on both dataset curation as well as design of predictive models. Here, we summarize the results of the SeqPIP competition whose objective was to develop comprehensive PPI predictive models from sequence information with high-quality bias-free interaction datasets. A training set of 2000 positive and 2000 negative interactions with sequences was given to each contestant. The methods were evaluated with three independent high-quality interaction test datasets.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.