Back

Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks

Mattsson, B.;Walters, W.

2026-06-30 Molecular Biology
10.64898/2026.06.29.735309 bioRxiv
Show abstract

Accurate prediction of protein-ligand binding affinity is a crucial goal in structure-based drug discovery, with the potential to significantly shorten development timelines. Recently, a new wave of machine learning models based on co-folding, such as Boltz-2 and IsoDDE, has demonstrated performance that matches or exceeds that of gold-standard physics-based methods like Free Energy Perturbation (FEP). This paper provides a critical assessment of these claims, revealing that current benchmarks are heavily influenced by data leakage, and proposes a new benchmark that explicitly controls for data leakage. We demonstrate that splitting by protein-sequence identity is inherently insufficient to prevent data leakage due to "target mirroring," in which homologous proteins with low overall sequence identity still exhibit highly correlated binding profiles. Our meta-analysis of documents in the ChEMBL 36 database identifies more than 6,000 such assay pairs and finds that leakage persists for sequence-identity thresholds as low as 0.2, well below the values commonly used in benchmarks today. Additionally, we show that a ligand-only baseline model, which lacks protein structural information, achieves surprisingly high performance on the FEP+ 4 and OpenFE benchmarks (r = 0.66 and r = 0.36, respectively). Our results indicate that current benchmarks tend to reward models for memorizing training data and exploiting localized leakage rather than truly learning biophysical principles. To address this issue, we propose the Novelty-Tiered Affinity Benchmark, in which the test data is partitioned into ligand novelty tiers. In the most challenging tier (Tanimoto similarity < 0.35), ligand-only models perform notably worse (r = 0.14), offering a clear baseline for evaluating genuine generalization. We argue that the field must move beyond sequence-based splits to ensure that AI-driven discovery translates into successful prospective laboratory research.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.2%
19.3%
2
Cell Systems
201 papers in training set
Top 0.2%
10.2%
3
Scientific Reports
3612 papers in training set
Top 8%
7.6%
4
PLOS Computational Biology
1863 papers in training set
Top 7%
5.1%
5
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 11%
4.6%
6
eLife
5828 papers in training set
Top 26%
4.5%
50% of probability mass above
7
Bioinformatics
1204 papers in training set
Top 5%
3.6%
8
Journal of Cheminformatics
29 papers in training set
Top 0.2%
3.4%
9
Patterns
78 papers in training set
Top 0.5%
3.4%
10
PLOS ONE
5266 papers in training set
Top 42%
2.5%
11
Nature Communications
5641 papers in training set
Top 39%
2.5%
12
Human Genetics and Genomics Advances
84 papers in training set
Top 1%
1.8%
13
Bioinformatics Advances
203 papers in training set
Top 3%
1.6%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.2%
15
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.2%
16
Journal of Molecular Biology
232 papers in training set
Top 2%
1.2%
17
Nature Methods
385 papers in training set
Top 5%
1.2%
18
Protein Science
246 papers in training set
Top 3%
1.2%
19
Communications Biology
993 papers in training set
Top 20%
1.2%
20
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.2%
1.0%
21
Cell Reports Methods
165 papers in training set
Top 3%
1.0%
22
Frontiers in Pharmacology
111 papers in training set
Top 3%
0.9%
23
International Journal of Molecular Sciences
494 papers in training set
Top 14%
0.9%
24
iScience
1154 papers in training set
Top 33%
0.9%
25
ACS Omega
105 papers in training set
Top 3%
0.9%
26
Nucleic Acids Research
1281 papers in training set
Top 13%
0.9%
27
Biophysical Journal
631 papers in training set
Top 4%
0.9%
28
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
29
mAbs
32 papers in training set
Top 0.5%
0.6%
30
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.6%