Back

PICKER-HG: a web server using random forests for classifying human genes into categories

Fabris, F.; Palmer, D.; Farooq, Z.; de Magalhaes, J. P.; Freitas, A. A.

2019-06-24 bioinformatics
10.1101/681460 bioRxiv
Show abstract

MotivationOne of the main challenges faced by biologists is how to extract valuable knowledge from the data produced by high-throughput genomic experiments. Although machine learning can be used for this, in general, machine learning tools on the web were not designed for biologist users. They require users to create suitable biological datasets and often produce results that are hard to interpret.\n\nObjectiveOur aim is to develop a freely available web server, named PerformIng Classification and Knowledge Extraction via Rules using random forests on Human Genes (PICKER-HG), aimed at biologists looking for a straightforward application of a powerful machine learning technique (random forests) to their data.\n\nResultsWe have developed the first web server that, as far as we know, dynamically constructs a classification dataset, given a list of human genes with annotations entered by the user, and outputs classification rules extracted of a Random Forest model. The web server can also classify a list of genes whose class labels are unknown, potentially assisting biologists investigating the association between class labels of interest and human genes.\n\nAvailabilityhttp://machine-learning-genomics.com/

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.