Binding Site-enhanced Sequence Pretraining and Out-of-cluster Meta-learning Predict Genome-Wide Chemical-Protein Interactions for Dark Proteins
Cai, T.; Xie, L.
Show abstract
Discovering chemical-protein interactions for millions of chemicals across the entire human and pathogen genomes is instrumental for chemical genomics, protein function prediction, drug discovery, and other applications. However, more than 90% of gene families remain dark, i.e., their small molecular ligands are undiscovered due to experimental limitations and human biases. Existing computational approaches typically fail when the unlabeled dark protein of interest differs from those with known ligands or structures. To address this challenge, we developed a deep learning framework PortalCG. PortalCG consists of four novel components: (i) a 3-dimensional ligand binding site enhanced sequence pre-training strategy to represent the whole universe of protein sequences in recognition of evolutionary linkage of ligand binding sites across gene families, (ii) an end-to-end pretraining-fine-tuning strategy to simulate the folding process of protein-ligand interactions and reduce the impact of inaccuracy of predicted structures on function predictions under a sequence-structure-function paradigm, (iii) a new out-of-cluster meta-learning algorithm that extracts and accumulates information learned from predicting ligands of distinct gene families (meta-data) and applies the meta-data to a dark gene family, and (iv) stress model selection that uses different gene families in the test data from those in the training and development data sets to facilitate model deployment in a real-world scenario. In extensive and rigorous benchmark experiments, PortalCG considerably outperformed state-of-the-art techniques of machine learning and protein-ligand docking when applied to dark gene families, and demonstrated its generalization power for off-target predictions and compound screenings under out-of-distribution (OOD) scenarios. Furthermore, in an external validation for the multi-target compound screening, the performance of PortalCG surpassed the human design. Our results also suggested that a differentiable sequence-structure-function deep learning framework where protein structure information serve as an intermediate layer could be superior to conventional methodology where the use of predicted protein structures for predicting protein functions from sequences. We applied PortalCG to two case studies to exemplify its potential in drug discovery: designing selective dual-antagonists of Dopamine receptors for the treatment of Opioid Use Disorder, and illuminating the undruggable human genome for targeting diseases that do not have effective and safe therapeutics. Our results suggested that PortalCG is a viable solution to the OOD problem in exploring the understudied protein functional space. Author SummaryMany complex diseases such as Alzheimers disease, mental disorders, and substance use disorders do not have effective and safe therapeutics due to the polygenic nature of diseases and the lack of thoroughly validate drug targets and their ligands. Identifying small molecule ligands for all proteins encoded in the human genome will provide new opportunity for drug discovery of currently untreatable diseases. However, the small molecule ligand of more than 90% gene families is completely unknown. Existing protein-ligand docking and machine learning methods often fail when the protein of interest is dissimilar to those with known functions or structures. We develop a new deep learning framework PortalCG for efficiently and accurately predicting ligands of understudied proteins which are out of reach of existing methods. Our method achieves unprecedented accuracy over state-of-the-arts by incorporating ligand binding site information and sequence-to-structure-to-function paradigm into a novel deep meta-learning algorithms. In a case study, the performance of PortalCG surpassed the human design. The proposed computational framework will shed new light into how chemicals modulate biological system as demonstrated by applications to drug repurposing and designing polypharmacology. It will open a new door to developing effective and safe therapeutics for currently incurable diseases. PortalCG can be extended to other scientific inquiries such as predicting protein-protein interactions and protein-nucleic acid recognition.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Attention-based approach to predict drug-target interactions across seven target superfamilies 98%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 97%
- Prediction of bacterial protein-compound interactions with only positive samples 97%
Similar papers in this journal
- Controlling astrocyte-mediated synaptic pruning signals for schizophrenia drug repurposing with Deep Graph Networks 96%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 96%
- Novel, provable algorithms for efficient ensemble-based computational protein design and their application to the redesign of the c-Raf-RBD:KRas protein-protein interface 95%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- A Graph-Attention-Based Deep Learning Network for Predicting Biotech-Small-Molecule Drug Interactions 96%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.