How Bias Shapes the Leaderboard: Scoring Function Performance Under Scrutiny
Graber, D.; Kopko, J.; Stockinger, P.; Nakandalage, R.; Kuhn, B.; Mishra, S.
Show abstract
Structure-based scoring functions leveraging machine learning have recently demonstrated superior performance over classical scoring functions, particularly on virtual screening benchmarks. However, due to the fundamental differences between their underlying model principles and architectures, it remains unclear to what extent performance stems from an understanding of molecular binding or from exploitation of systemic biases. Thus, disentangling the factors underlying benchmark performance is essential for determining whether a scoring function will generalize to novel chemical space and succeed in prospective drug discovery. To address this need, we present a case study investigating the nature and impact of systemic biases on benchmark comparisons between different scoring function paradigms. By systematically analyzing the evaluation workflows of prominent models, we reveal pocket bias, a form of spatial coordinate frame leakage arising from static binding pocket extraction, which artificially inflates benchmark performance. To progressively eliminate these sources of bias, we benchmarked two selected graph neural network scoring functions against two minmalist machine learning models and a classical scoring function under four increasingly stringent evaluation levels, successively removing pocket bias, reducing structural data leakage, and finally evaluating on out-of-distribution (OOD) protein targets. Upon removal of pocket bias and structural data leakage, the performance of all machine learning models dropped substantially. When evaluated on out-of-distribution protein families, the classical baseline AutoDock Vina outperformed the machine learning models in five of seven virtual screening tasks and dominated the docking power evaluation. Our findings indicate that benchmark performance can be heavily shaped by evaluation design and dataset artifacts, potentially overshadowing algorithmic improvements. While the tested machine learning models remain heavily dependent on encountering familiar data distributions to achieve competitive results, AutoDock Vina demonstrated superior generalization capacity on OOD targets. This work underscores the critical need for rigorous, artifact-free benchmarking protocols to guide the development of truly prospective machine learning models for virtual screening.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 97%
- DENVIS: scalable and high-throughput virtual screening using graph neural networks with atomic and surface protein pocket features 97%
- Dataset Augmentation Allows Deep Learning-Based Virtual Screening To Better Generalize To Unseen Target Classes, And Highlight Important Binding Interactions 97%
Similar papers in this journal
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 96%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 96%
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.