Infinite Physical Monkey: Do Deep Learning Methods Really Perform Better in Conformation Generation?
Zhang, H.; Zhang, J.; Zhao, H.; Jiang, D.; Deng, Y.
Show abstract
Conformation Generation is a fundamental problem in drug discovery and cheminformatics. Generally, it can be categorized into three different classes according to physical scales, i.e., micro molecule (organic), meso molecule (nano particle-like), and macro molecule (protein and nucleic acid). Organic molecule conformation generation, particularly in vacuum and protein pocket environments, is most relevant to drug design. Recently, with the development of geometric neural networks, the data-driven schemes have been successfully applied in this field, both for molecular conformation generation (in vacuum) and binding pose generation (in protein pocket). The former beats the traditional ETKDG method, while the latter achieves similar accuracy compared with the widely used molecular docking software. Although these methods have shown promising results for real-world drug design campaigns, some researchers have recently questioned whether deep learning (DL) methods perform better in molecular conformation generation via a "parameter-free" method. To our surprise, what they have designed is some kind analogous to the famous infinite monkey theorem, the monkeys that are even equipped with physics education. To discuss the feasibility of their proving, we constructed a real infinite stochastic monkey for molecular conformation generation, showing that even with a more stochastic sampler for geometry generation, the coverage of the benchmark QM-computed conformations are higher than those of most DL-based methods. By extending their physical monkey algorithm for binding pose prediction (with 2000 random samples), we also discover that the successful docking rate also achieves near-best performance among existing DL-based docking models. Thus, though their conclusions are right, their proof process needs more concern. In addition to evaluating the rationality of their algorithms and conclusions, we dig into the inspirations of infinite physical monkeys. We find that, for docking pose generation, DL-based models truly learn the interaction rules between residues and ligands, and discover an inductive bias hidden in the training of the pocket-given docking problem. The code of the proposed algorithm could be found at: https://github.com/HaotianZhangAI4Science/infinite-physical-monkey. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/531607v2_ufig1.gif" ALT="Figure 1"> View larger version (126K): org.highwire.dtl.DTLVardef@188ac2eorg.highwire.dtl.DTLVardef@1e006f3org.highwire.dtl.DTLVardef@e83f8forg.highwire.dtl.DTLVardef@1a4c7cd_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOTOC:C_FLOATNO Infinite Physical Monkey. This image was created with the assistance of DALL{middle dot}E 2 C_FIG
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Rationalize the Functional Roles of Protein-Protein Interactions in Targeted Protein Degradation by Kinetic Monte-Carlo Simulations 96%
- Predicting the Activities of Drug Excipients on Biological Targets using One-Shot Learning 95%
- 3D-Scaffold: Deep Learning Framework to Generate 3D Coordinates of Drug-like Molecules with Desired Scaffolds. 95%
Similar papers in this journal
- Protein Domain-Based Prediction of Compound-Target Interactions and Experimental Validation on LIM Kinases 93%
- TranSynergy: Mechanism-Driven Interpretable Deep Neural Network for the Synergistic Prediction and Pathway Deconvolution of Drug Combinations 93%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 93%
Similar papers in this journal
- ANUBI: A Platform for Affinity Optimization of Proteins and Peptides in Drug Design 94%
- Ligand Gaussian accelerated molecular dynamics 2 (LiGaMD2): Improved calculations of ligand binding thermodynamics and kinetics with closed protein pocket 94%
- Flexible Fitting of Biomolecular Structures to Atomic Force Microscopy Images via Biased Molecular Simulations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.