Are under-studied proteins under-represented? How to fairly evaluate link prediction algorithms in network biology
Yilmaz, S.; Yorgancioglu, K.; Koyuturk, M.
Show abstract
For biomedical applications, new link prediction algorithms are continuously being developed and these algorithms are typically evaluated computationally, using test sets generated by sampling the edges uniformly at random. However, as we demonstrate, this evaluation approach introduces a bias towards "rich nodes", i.e., those with higher degrees in the network. More concerningly, this bias persists even when different network snapshots are used for evaluation, as recommended in the machine learning community. This creates a cycle in research where newly developed algorithms generate more knowledge on well-studied biological entities while under-studied entities are commonly overlooked. To overcome this issue, we propose a weighted validation setting specifically focusing on under-studied entities and present AWARE strategies to facilitate bias-aware training and evaluation of link prediction algorithms. These strategies can help researchers gain better insights from computational evaluations and promote the development of new algorithms focusing on novel findings and under-studied proteins. TeaserSystematically characterizes and mitigates bias toward well-studied proteins in the evaluation pipeline for machine learning. Code and data availabilityAll materials (code and data) to reproduce the analyses and figures in the paper is available in figshare (doi:10.6084/m9.figshare.21330429). The code for the evaluation framework implementing the proposed strategies is available at github{dagger}. We provide a web tool{ddagger} to assess the bias in benchmarking data and to generate bias-adjusted test sets.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.