Back

An ensemble method for predicting and designing of druggable proteins.

Jain, S.; Gupta, S.; Raghava, G. P. S.

2024-10-18 bioinformatics
10.1101/2024.10.17.618792 bioRxiv
Show abstract

In the past, numerous proteins/peptides have been discovered, which have a wide range of therapeutic properties like anticancer, antimicrobial, antihypertensive. Only few hundreds of proteins are druggable (approved by US FDA), most of the proteins fails in clinical trials. In this study, an attempt had been made to understand properties of FDA approved proteins to develop models for predicting druggable proteins. Our main dataset 356 FDA approved proteins as positive dataset and equal number of randomly selected proteins as negative dataset. We used 80% data for training and 20% for independent validation, no protein in validation dataset have more than 40% similarity with any protein in training dataset. We deployed machine learning based models using a five-fold cross-validation and test on validation dataset. Our random forest-based model developed using SVC-L1 selected features obtained maximum performance AUC of 0.80 with MCC 0.61 on validation data. In addition to this, we performed MERCI-based motif analysis to find motifs in druggable proteins. Finally, we developed an ensemble-based method combining best performing machine learning model with motifs and achieved AUC 0.92 with MCC 0.83 on independent validation dataset. We developed a web server and standalone package ThPPred to facilitate scientific community in predicting and designing druggable proteins (https://webs.iiitd.edu.in/raghava/thppred/). HighlightsO_LIAnalysis of FDA approved or druggable proteins C_LIO_LIDiscrimination of druggable and non-druggable proteins C_LIO_LIMachine learning based models for predicting druggable molecules C_LIO_LIIdentification of motifs in druggable proteins C_LIO_LIA web server for providing service to community C_LI

Published in PROTEOMICS – Clinical Applications · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.