Back

Performance Attribution in pLM-Based Biological Relation Prediction

Zhu, K.; Zhao, W.; Zhang, Y.; Xia, Z.

2026-07-02 bioinformatics
10.64898/2026.06.28.735130 bioRxiv
Show abstract

Protein language models (pLMs) have enabled strong benchmark performance in biological relation-prediction tasks, but aggregate metrics do not identify which sources of information support that performance. We examined three case studies-MetaESI, DeepGNHV, and SAGEPhos-using frozen pLM full-input baselines, endpoint- or site-restricted controls, clean train-derived prior controls, and a restriction-matched selector control. Frozen pLM-derived inputs coupled to generic downstream learners reached AUROC/AUPRC of 0.827/0.703 for the current MetaESI full-input rerun, 0.922/0.708 for a DeepGNHV two-endpoint baseline, and 0.896/0.893 for SAGEPhos. Restricted controls retained task-dependent signal. Most notably, a self-label-excluding MetaESI endpoint-frequency control reached 0.846/0.676 under the row split, numerically close to the frozen ESM2 full-input reference despite using no sequence embeddings. Clean full-catalog DeepGNHV endpoint priors and SAGEPhos kinase/substrate/site-window priors provided supplementary diagnostics rather than direct architecture-contribution estimates. In a separate fixed frozen-pooling LightGBM experiment, GARD-selected pooling did not show higher observed performance than count-matched random token pooling. Endpoint-cold diagnostics showed performance degradation, while train-label shuffling returned discrimination to approximately chance; hard- or matched-negative and family- or homology-aware evaluations were not available across the case studies. These findings do not diagnose leakage, imply memorization, exclude biological learning, or invalidate the evaluated models. Rather, they show that benchmark utility and performance attribution are distinct: architecture-specific, relation-specific, selector-specific, and generalization claims require controls matched to the interpretation being made.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Briefings in Bioinformatics
354 papers in training set
Top 0.1%
26.1%
2
Scientific Reports
3612 papers in training set
Top 6%
8.8%
3
Nature Communications
5641 papers in training set
Top 21%
7.8%
4
Bioinformatics
1204 papers in training set
Top 4%
6.2%
5
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.8%
5.1%
50% of probability mass above
6
Cell Systems
201 papers in training set
Top 1%
3.2%
7
Nature Machine Intelligence
70 papers in training set
Top 0.9%
3.2%
8
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.4%
9
Frontiers in Bioinformatics
49 papers in training set
Top 0.3%
2.4%
10
Molecular Systems Biology
162 papers in training set
Top 1%
1.7%
11
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 28%
1.7%
12
PLOS Computational Biology
1863 papers in training set
Top 15%
1.7%
13
New Phytologist
346 papers in training set
Top 4%
1.7%
14
Communications Biology
993 papers in training set
Top 17%
1.5%
15
Nature Methods
385 papers in training set
Top 5%
1.5%
16
Genome Biology
637 papers in training set
Top 6%
1.5%
17
PLOS ONE
5266 papers in training set
Top 53%
1.3%
18
iScience
1154 papers in training set
Top 23%
1.3%
19
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
20
eLife
5828 papers in training set
Top 58%
1.1%
21
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.1%
22
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
23
Protein Science
246 papers in training set
Top 3%
1.0%
24
Patterns
78 papers in training set
Top 2%
1.0%
25
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
26
GigaScience
212 papers in training set
Top 5%
0.8%
27
Nature
645 papers in training set
Top 12%
0.6%
28
Frontiers in Genetics
230 papers in training set
Top 7%
0.6%