LEP-AD: Language Embedding of Proteins and Attention to Drugs predicts drug target interactions
Daga, A.; Khan, S. A.; Cabrero, D. G.; Hoehndorf, R.; Kiani, N.; Tegner, J.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWPredicting drug-target interactions is a tremendous challenge for drug development and lead optimization. Recent advances include training algorithms to learn drug-target interactions from data and molecular simulations. Here we utilize Evolutionary Scale Modeling (ESM-2) models to establish a Transformer protein language model for drug-target interaction predictions. Our architecture, LEP-AD, combines pre-trained ESM-2 and Transformer-GCN models predicting binding affinity values. We report new best-in-class state-of-the-art results compared to competing methods such as SimBoost, DeepCPI, Attention-DTA, GraphDTA, and more using multiple datasets, including Davis, KIBA, DTC, Metz, ToxCast, and STITCH. Finally, we find that a pre-trained model with embedding of proteins (the LED-AD) outperforms a model using an explicit alpha-fold 3D representation of proteins (e.g., LEP-AD supervised by Alphafold). The LEP-AD model scales favorably in performance with the size of training data. Code available at https://github.com/adaga06/LEP-AD
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 96%
- Deep learning integration of molecular and interactome data for protein-compound interaction prediction 94%
- Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map 94%
Similar papers in this journal
Similar papers in this journal
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 95%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 95%
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 95%
Similar papers in this journal
- Retro Drug Design: From Target Properties to Molecular Structures 97%
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 95%
- Streamlining Computational Fragment-Based Drug Discovery through Evolutionary Optimization Informed by Ligand-Based Virtual Prescreening 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.