Node-degree aware edge sampling mitigates inflated classification performance in biomedical graph representation learning
Cappelletti, L.; Rekerle, L.; Fontana, T.; Hansen, P.; Casiraghi, E.; Ravanmehr, V.; Mungall, C. J.; Yang, J.; Spranger, L.; Karlebach, G.; Caufield, J. H.; Carmody, L.; Coleman, B.; Oprea, T.; Reese, J.; Valentini, G.; Robinson, P. N.
Show abstract
Graph representation learning is a family of related approaches that learn low-dimensional vector representations of nodes and other graph elements called embeddings. Embeddings approximate characteristics of the graph and can be used for a variety of machine-learning tasks such as novel edge prediction. For many biomedical applications, partial knowledge exists about positive edges that represent relationships between pairs of entities, but little to no knowledge is available about negative edges that represent the explicit lack of a relationship between two nodes. For this reason, classification procedures are forced to assume that the vast majority of unlabeled edges are negative. Existing approaches to sampling negative edges for training and evaluating classifiers do so by uniformly sampling pairs of nodes. We show here that this sampling strategy typically leads to sets of positive and negative edges with imbalanced edge degree distributions. Using representative homogeneous and heterogeneous biomedical knowledge graphs, we show that this strategy artificially inflates measured classification performance. We present a degree-aware node sampling approach for sampling negative edge examples that mitigates this effect and is simple to implement.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Connectivity Measures for Signaling Pathway Topologies 96%
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 95%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.