Back

Shape-restrained modelling of protein-small molecule complexes with HADDOCK

Koukos, P. I.; Reau, M. F.; Bonvin, A. M. J. J.

2021-06-10 bioinformatics
10.1101/2021.06.10.447890 bioRxiv
Show abstract

Small molecule docking remains one of the most valuable computational techniques for the structure prediction of protein-small molecule complexes. It allows us to study the interactions between compounds and the protein receptors they target at atomic detail, in a timely and efficient manner. Here we present a new protocol in HADDOCK, our integrative modelling platform, which incorporates homology information for both receptor and compounds. It makes use of HADDOCKs unique ability to integrate information in the simulation to drive it toward conformations which agree with the provided data. The focal point is the use of shape restraints derived from homologous compounds bound to the target receptors. We have developed two protocols: In the first, the shape is composed of fake atom beads based on the position of the heavy atoms of the homologous template compound, whereas in the second the shape is additionally annotated with pharmacophore data, for some or all beads. For both protocols, ambiguous distance restraints are subsequently defined between those beads and the heavy atoms of the ligand to be docked. We have benchmarked the performance of these protocols with a fully unbound version of the widely used DUD-E dataset. In this unbound docking scenario, our template/shape-based docking protocol reaches an overall success rate of 81% on 99 complexes, which is close to the best results reported for bound docking on the DUD-E dataset. Table of contents graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=97 SRC="FIGDIR/small/447890v1_ufig1.gif" ALT="Figure 1"> View larger version (26K): org.highwire.dtl.DTLVardef@15faf1dorg.highwire.dtl.DTLVardef@e1d311org.highwire.dtl.DTLVardef@1e8225eorg.highwire.dtl.DTLVardef@1285ed0_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.