Evaluation of Machine Learning Models for Aqueous Solubility Prediction in Drug Discovery
Liu, S.; Zhang, Y.; Xue, N.
Show abstract
Determining the aqueous solubility of the chemical compound is of great importance in-silico drug discovery. However, correctly and rapidly predicting the aqueous solubility remains a challenging task. This paper explores and evaluates the predictability of multiple machine learning models in the aqueous solubility of compounds. Specifically, we apply a series of machine learning algorithms, including Random Forest, XG-Boost, LightGBM, and CatBoost, on a well-established aqueous solubility dataset (i. e., the Huuskonen dataset) of over 1200 compounds. Experimental results show that even traditional machine learning algorithms can achieve satisfactory performance with high accuracy. In addition, our investigation goes beyond mere prediction accuracy, delving into the interpretability of models to identify key features and understand the molecular properties that influence the predicted outcomes. This study sheds light on the ability to use machine learning approaches to predict compound solubility, significantly shortening the time that researchers spend on new drug discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning based predictive modeling to screen natural compounds against TNF-alpha for the potential management of Rheumatoid Arthritis: Virtual screening to comprehensive in silico investigation 95%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- PharmaNet: Pharmaceutical discovery with deep recurrent neural networks. 95%
Similar papers in this journal
- Advancements in Ligand-Based Virtual Screening through the Synergistic Integration of Graph Neural Networks and Expert-Crafted Descriptors 95%
- Streamlining Computational Fragment-Based Drug Discovery through Evolutionary Optimization Informed by Ligand-Based Virtual Prescreening 95%
- Retro Drug Design: From Target Properties to Molecular Structures 94%
Similar papers in this journal
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 94%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 94%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 93%
Similar papers in this journal
- Deep Learning Modelling of Androgen Receptor Responses to Prostate Cancer Therapies 94%
- SSnet: A Deep Learning Approach for Protein-Ligand Interaction Prediction 93%
- Effect of Delta and Omicron mutations on the RBD-SD1 do-main of the Spike protein in SARS-CoV-2 and the Omicron mutations on RBD-ACE2 interface complex 91%
Similar papers in this journal
- Designing of thermostable proteins with a desired melting temperature 93%
- DeepLPI: a novel deep learning-based model for protein-ligand interaction prediction for drug repurposing 92%
- Smart Distributed Data Factory: Volunteer Computing Platform for Active Learning-Driven Molecular Data Acquisition 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.