Exploration of Chemical Space with Partial Labeled Noisy Student Self-Training for Improving Deep Learning: Application to Drug Metabolism
Liu, Y.; Lim, H.; Xie, L.
Show abstract
MotivationDrug discovery is time-consuming and costly. Machine learning, especially deep learning, shows a great potential in accelerating the drug discovery process and reducing its cost. A big challenge in developing robust and generalizable deep learning models for drug design is the lack of a large amount of data with high quality and balanced labels. To address this challenge, we developed a self-training method PLANS that exploits millions of unlabeled chemical compounds as well as partially labeled pharmacological data to improve the performance of neural network models. ResultWe evaluated the self-training with PLANS for Cytochrome P450 binding activity prediction task, and proved that our method could significantly improve the performance of the neural network model with a large margin. Compared with the baseline deep neural network model, the PLANS-trained neural network model improved accuracy, precision, recall, and F1 score by 13.4%, 12.5%, 8.3%, and 10.3%, respectively. The self-training with PLANS is model agnostic, and can be applied to any deep learning architectures. Thus, PLANS provides a general solution to utilize unlabeled and partially labeled data to improve the predictive modeling for drug discovery. AvailabilityThe code that implements PLANS is available at https://github.com/XieResearchGroup/PLANS
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Data Imbalance in Drug Response Prediction - Multi-Objective Optimization Approach in Deep Learning Setting 97%
- Kinase-Inhibitor Binding Affinity Prediction with Pretrained Graph Encoder and Language Model 96%
- DeepDDS: deep graph neural network with attention mechanism to predict synergistic drug combinations 96%
Similar papers in this journal
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 96%
- Deep learning integration of molecular and interactome data for protein-compound interaction prediction 95%
- Evaluation of network architecture and data augmentation methods for deep learning in chemogenomics 95%
Similar papers in this journal
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 94%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 93%
- Predicting the pathway involvement of metabolites annotated in the MetaCyc knowledgebase 93%
Similar papers in this journal
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 95%
- FLONE: fully Lorentz network embedding for inferring novel drug targets 95%
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.