Back

Are under-studied proteins under-represented? How to fairly evaluate link prediction algorithms in network biology

Yilmaz, S.; Yorgancioglu, K.; Koyuturk, M.

2022-10-17 systems biology
10.1101/2022.10.13.511953 bioRxiv
Show abstract

For biomedical applications, new link prediction algorithms are continuously being developed and these algorithms are typically evaluated computationally, using test sets generated by sampling the edges uniformly at random. However, as we demonstrate, this evaluation approach introduces a bias towards "rich nodes", i.e., those with higher degrees in the network. More concerningly, this bias persists even when different network snapshots are used for evaluation, as recommended in the machine learning community. This creates a cycle in research where newly developed algorithms generate more knowledge on well-studied biological entities while under-studied entities are commonly overlooked. To overcome this issue, we propose a weighted validation setting specifically focusing on under-studied entities and present AWARE strategies to facilitate bias-aware training and evaluation of link prediction algorithms. These strategies can help researchers gain better insights from computational evaluations and promote the development of new algorithms focusing on novel findings and under-studied proteins. TeaserSystematically characterizes and mitigates bias toward well-studied proteins in the evaluation pipeline for machine learning. Code and data availabilityAll materials (code and data) to reproduce the analyses and figures in the paper is available in figshare (doi:10.6084/m9.figshare.21330429). The code for the evaluation framework implementing the proposed strategies is available at github{dagger}. We provide a web tool{ddagger} to assess the bias in benchmarking data and to generate bias-adjusted test sets.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.9%
22.6%
2
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.1%
12.8%
3
PLOS Computational Biology
1863 papers in training set
Top 3%
9.8%
4
Bioinformatics Advances
203 papers in training set
Top 0.4%
8.0%
50% of probability mass above
5
Scientific Reports
3612 papers in training set
Top 16%
5.5%
6
BMC Bioinformatics
457 papers in training set
Top 2%
4.1%
7
Patterns
78 papers in training set
Top 0.4%
4.1%
8
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.3%
9
PLOS ONE
5266 papers in training set
Top 38%
3.2%
10
Frontiers in Systems Biology
10 papers in training set
Top 0.1%
2.4%
11
Entropy
21 papers in training set
Top 0.1%
1.9%
12
iScience
1154 papers in training set
Top 16%
1.7%
13
GigaScience
212 papers in training set
Top 3%
1.4%
14
Biomolecules
100 papers in training set
Top 2%
1.1%
15
npj Systems Biology and Applications
125 papers in training set
Top 1%
1.1%
16
Nature Communications
5641 papers in training set
Top 50%
1.1%
17
Frontiers in Bioinformatics
49 papers in training set
Top 1%
1.0%
18
IEEE Access
35 papers in training set
Top 1%
1.0%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 41%
0.9%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.9%
21
Frontiers in Genetics
230 papers in training set
Top 6%
0.6%
22
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.6%
23
Computers in Biology and Medicine
128 papers in training set
Top 5%
0.6%