Structural Context Determines Docking Engine Performance: A Family-Stratified Benchmark of Six Engines
Alejo, K.; Fisher, S.; Kalluri, T.; More, B.; Rajgure, H.; Panda, P. K.; Korban, C.; Chung, C.
Show abstract
Molecular docking and co-folding engines are widely used to prioritize compounds for wet-lab validation, yet their accuracy is known to vary substantially across protein targets for reasons that remain only qualitatively understood. Here we benchmark six docking and co-folding engines (RevDock, DiffDock, Boltz2, AutoDock-GPU, rDock, and PandaDock) across 14 protein families, evaluating scoring power, ranking power, docking power, and physical validity. Rather than treating engine performance as protein-family-specific, we classify all 14 families into six mechanistic groups according to which of four scoring-function simplifications, rigid receptor, pairwise additivity, fixed point charges, and implicit solvent, is most severely stressed by that familys binding site. This framework helps explain, rather than simply describe, where each engine succeeds or fails: RevDocks CNN rescoring layer mitigates the pairwise additivity and fixed-charge limitations relative to physics-only scoring, achieving the highest overall pose accuracy (73.3% of poses [≤] 2.0 [A] RMSD), while Boltz2s sequence-based co-folding bypasses the rigid-receptor assumption and achieves comparable affinity correlation (mean Pearson r {approx} 0.60 for both engines). PandaDock, run with expanded conformational sampling, matches RevDock on pose accuracy (72.1% of poses [≤] 2.0 [A], lowest median RMSD at 0.96 [A]) and exceeds AutoDock-GPU on affinity correlation (mean r = 0.460), indicating that the performance of a physics-based scoring function is limited as much by search adequacy as by the scoring function itself. These results suggest that engine selection for a docking or co-folding campaign should be guided less by an engines aggregate benchmark ranking and more by which of these four structural and physical characteristics dominate the target of interest.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
- CANDOCK: Chemical atomic network based hierarchical flexible docking algorithm using generalized statistical potentials 96%
- Prioritizing virtual screening with interpretable interaction fingerprints 96%
Similar papers in this journal
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 94%
- Large-scale Annotation of Biochemically Relevant Pockets and Tunnels in Cognate Enzyme-Ligand Complexes 94%
- Deep learning molecular interaction motifs from Receptor structure alone 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.