LigEGFR: Spatial graph embedding and molecular descriptors assisted bioactivity prediction of ligand molecules for epidermal growth factor receptor on a cell line-based dataset
Virakarin, P.; Saengnil, N.; Boonyarit, B.; Kinchagawat, J.; Laotaew, R.; Saeteng, T.; Nilsu, T.; Suvannang, N.; Rungrotmongkol, T.; Nutanong, S.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSMotivationC_ST_ABSLung cancer is a chronic non-communicable disease and is the cancer with the worlds highest incidence in the 21st century. One of the leading mechanisms underlying the development of lung cancer in nonsmokers is an amplification of the epidermal growth factor receptor (EGFR) gene. However, laboratories employing conventional processes of drug discovery and development for such targets encounter several pain-points that are cost- and time-consuming. Moreover, high failure rates are caused by efficacy and safety problems during research and development. Therefore, it is imperative to develop improved methods for drug discovery. Herein, we developed a deep learning model with spatial graph embedding and molecular descriptors based on predicting pIC50 potency estimates of small molecules and classifying hit compounds against the human epidermal growth factor receptor (LigEGFR). The model was generated with a large-scale cell line-based dataset containing broad lists of chemical features. ResultsLigEGFR outperformed baseline machine learning models for predicting pIC50. Our model was notable for higher performance in hit compound classification, compared to molecular docking and machine learning approaches. The proposed predictive model provides a powerful strategy that potentially helps researchers overcome major challenges in drug discovery and development processes, leading to a reduction of failure to discover novel hit compounds. AvailabilityWe provide an online prediction platform and the source code that are freely available at https://ligegfr.vistec.ist, and https://github.com/scads-biochem/LigEGFR, respectively. Key pointsO_LILigEGFR is a regression model for predicting pIC50 that was developed for the human EGFR target. It can also be applied to hit compound classification (pIC50 [≥] 6) and has a higher performance than baseline machine learning algorithms and molecular docking approaches. C_LIO_LIOur spatial graph embedding and molecular descriptors based approach notably exhibited a high performance in predicting pIC50 of small molecules against human EGFR. C_LIO_LINon-hashed and hashed molecular descriptors were revealed to have the highest predictive performance by using in a convolutional layers and a fully connected layers, respectively. C_LIO_LIOur model used a large-scale and non-redundant dataset to enhance the diversity of the small molecules. The model showed robustness and reliability, which was evaluated by y-randomization and applicability domain analysis (ADAN), respectively. C_LIO_LIWe developed a user-friendly online platform to predict pIC50 of small molecules and classify the hit compounds for the drug discovery process of the EGFR target. C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fragment Linker Prediction Using Deep Encoder-Decoder Network for PROTAC Drug Design 98%
- Benchmarking of Small Molecule Feature Representations for hERG, Nav1.5, and Cav1.2 Cardiotoxicity Prediction 97%
- Deep Learning-based Ligand Design using Shared Latent Implicit Fingerprints from Collaborative Filtering 97%
Similar papers in this journal
Similar papers in this journal
- In Silico Identification of Potential Inhibitors of Mycobacterium tuberculosis DNA Gyrase from Phytoconstituents of Indian Medicinal Plants 96%
- Combining Multi-Dimensional Molecular Fingerprints to Predict hERG Cardiotoxicity of Compounds 94%
- ChAlPred: A Web Server for Prediction of Allergenicity of Chemical Compounds 94%
Similar papers in this journal
- Comprehensive machine learning boosts structure-based virtual screening for PARP1 inhibitors 96%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 96%
- Deep learning integration of molecular and interactome data for protein-compound interaction prediction 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.