Back

ProteinDock: A physics-informed layer to improve protein-protein docking reliability

Rajagopal, G.; Spina, S. C.; Bailey, J. S.; Kimmel, B. R.

2026-07-20 bioengineering
10.64898/2026.07.17.739238 bioRxiv
Show abstract

Computational modeling provides geometric insight into protein-protein interactions without requiring the resources of experimentation. However, reliability can be hindered when modeling proteins with distinctive features, such as antibodies, that use flexible, polar-rich loops to bind antigens. We developed ProteinDock, a physics-based tool that can be used in combination with leading modeling programs to improve the reliability of protein-protein docking; this work provides a case study of antibody-antigen interfaces. ProteinDock was layered onto Rosetta for docking unbound experimentally determined structures, and when evaluated on Docking Benchmark Set 5.5, generated CAPRI acceptable-quality or better for 80.2% of targets, an improvement of 32.8 percentage points over vanilla Rosettas 47.4% on the same dataset. To improve protein-protein prediction reliability from sequence inputs, we demonstrate that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools. We show that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions. A graphical user interface for layering ProteinDock has been created and is available at https://github.com/Kimmel-Lab/proteindock and https://proteindock.com/. TOC Figure O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=103 SRC="FIGDIR/small/739238v1_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@45183dorg.highwire.dtl.DTLVardef@3a7e2borg.highwire.dtl.DTLVardef@31664dorg.highwire.dtl.DTLVardef@1337eac_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.