Back

Optimizing InterProScan feature processing generates a surprisingly good protein function prediction method

Tiittanen, H.; Holm, L.; Toronen, P.

2022-08-26 bioinformatics
10.1101/2022.08.10.503467 bioRxiv
Show abstract

MotivationAutomated protein Function Prediction (AFP) is an intensively studied topic. Most of this research focuses on methods that combine multiple data sources, while fewer articles look for the most efficient ways to use a single data source. Therefore, we wanted to test how different preprocessing methods and classifiers would perform in the AFP task when we process the output from the InterProscan (IPS). Especially, we present novel preprocessing methods, less used classifiers and inclusion of species taxonomy. We also test classifier stacking for combining tested classifier results. Methods are tested with in-house data and CAFA3 competition evaluation data. ResultsWe show that including IPS localisation and taxonomy to the data improves results. Also the stacking improves the performance. Surprisingly, our best performing methods outperformed all international CAFA3 competition participants in most tests. Altogether, the results show how preprocessing and classifier combinations are beneficial in the AFP task. Contactpetri.toronen(AT)helsinki.fi Supplementary informationSupplementary text is available at the project web site http://ekhidna2.biocenter.helsinki.fi/AFP/ and at the end of this document.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.