Overcoming the accuracy-generalization tradeoff in docking and scoring for prospective virtual screening
Petrosyan, G.; Altunyan, V.; Ghukasyan, T.; Abramyan, T. M.; Arakelov, G.; Davtyan, A.; Aghajanyan, T.; Nakipov, I.; Navasardyan, G.; Fahradyan, A.; Tunanyan, H.; Simonyan, A.; Janczyk, P. Ł; De Silva, D.; Saribekyan, H.; Tsidilkovski, L.; Arakelov, V.; Ginoyan, N.; Mnatsakanyan, H.; Ratnikov, M.; Smbatyan, K.; Papoyan, A.; Papoian, G. A.
Show abstract
Virtual screening promises access to tens of billions of synthetically accessible, diverse compounds, yet it is rarely used as a primary hit-discovery strategy in contemporary drug-discovery campaigns. We argue that this gap reflects the real-world underperformance of the underlying docking and scoring methods: classical docking is generalizable but limited in accuracy by simple functional forms and parsimonious parameterization, whereas recent machine-learning approaches are highly expressive but do not generalize well to novel molecules and pockets, their reported accuracy often inflated by train-test leakage. To address these challenges, we introduce DODock and DOScore, docking and scoring ML/physics hybrid frameworks that are also highly expressive, yet generalize much better out of distribution compared with the prior ML approaches. This generalization has been prospectively tested in several ways. First, DODocks blind prediction of a drug candidate binding to PCSK9 was compared to the crystal structure that was subsequently solved, recovering the binding pose to 1.2 [A] RMSD. We also used DODock and DOScore in prospective virtual screening campaigns against four therapeutic targets, spanning an ectoenzyme (CD73), a kinase (IRAK4), an extended-substrate protease (FXI), and an allosteric inhibition of protein-protein interface (IL17). These screens yielded many chemically novel, biochemically and cellularly active inhibitors. In the case of CD73, which is a historically challenging target for virtual screening, our screen resulted in a roughly hundredfold improvement in hit rate over a recent machine-learning screen. Our results indicate that the apparent ceiling in the accuracy of virtual screening that seemed to have somewhat plateaued in the last two decades is not fundamental, and that structure-based interrogation of ultralarge chemical space may eventually become a credible primary route to novel chemical matter.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Transfer learning enables discovery of sub-micromolar antibacterials for ESKAPE pathogens from ultra-large chemical spaces 95%
- Deep Generative Design with 3D Pharmacophoric Constraints 95%
- Structure-Aware Dual-Target Drug Design through Collaborative Learning of Pharmacophore Combination and Molecular Simulation 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.