An ensemble method for predicting and designing of druggable proteins.
Jain, S.; Gupta, S.; Raghava, G. P. S.
Show abstract
In the past, numerous proteins/peptides have been discovered, which have a wide range of therapeutic properties like anticancer, antimicrobial, antihypertensive. Only few hundreds of proteins are druggable (approved by US FDA), most of the proteins fails in clinical trials. In this study, an attempt had been made to understand properties of FDA approved proteins to develop models for predicting druggable proteins. Our main dataset 356 FDA approved proteins as positive dataset and equal number of randomly selected proteins as negative dataset. We used 80% data for training and 20% for independent validation, no protein in validation dataset have more than 40% similarity with any protein in training dataset. We deployed machine learning based models using a five-fold cross-validation and test on validation dataset. Our random forest-based model developed using SVC-L1 selected features obtained maximum performance AUC of 0.80 with MCC 0.61 on validation data. In addition to this, we performed MERCI-based motif analysis to find motifs in druggable proteins. Finally, we developed an ensemble-based method combining best performing machine learning model with motifs and achieved AUC 0.92 with MCC 0.83 on independent validation dataset. We developed a web server and standalone package ThPPred to facilitate scientific community in predicting and designing druggable proteins (https://webs.iiitd.edu.in/raghava/thppred/). HighlightsO_LIAnalysis of FDA approved or druggable proteins C_LIO_LIDiscrimination of druggable and non-druggable proteins C_LIO_LIMachine learning based models for predicting druggable molecules C_LIO_LIIdentification of motifs in druggable proteins C_LIO_LIA web server for providing service to community C_LI
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Designing of a next generation multiepitope based vaccine (MEV) against SARS-COV-2: Immunoinformatics and in silico approaches 96%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- Identification of Natural Antiviral Drug Candidates Against Tilapia Lake Virus: Computational Drug Design Approaches 95%
Similar papers in this journal
Similar papers in this journal
- Molecular Glue-Design-Evaluator (MOLDE): An Advanced Method for In-Silico Molecular Glue Design 95%
- Utilizing Heteroatom Types and Numbers from Extensive Ligand Libraries to Develop Novel hERG Blocker QSAR Models Using Machine Learning-based Classifiers 94%
- Support Vector Machine based prediction models for drug repurposing and designing novel drugs for colorectal cancer 94%
Similar papers in this journal
- AntiCP 2.0: An updated model for predicting anticancer peptides 99%
- DBpred: A deep learning method for the prediction of DNA interacting residues in protein sequences 95%
- Design of an Epitope-Based Peptide Vaccine against the Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2): A Vaccine-informatics Approach 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.