Back

Contrastive Learning for Graph-Based Biological Interaction Discovery: Insights from Oncologic Pathways

Nguyen, P.-N.

2024-07-24 bioinformatics
10.1101/2024.07.23.604746 bioRxiv
Show abstract

BackgroundContrastive learning has emerged as a pivotal technique in representation learning, particularly for self-supervised and unsupervised tasks. Link prediction, crucial for network analysis, forecasts the formation of connections between nodes. Machine learning enhances link prediction by learning patterns from data, leading to improved performance and scalability. MethodIn this study, we propose a contrastive learning approach tailored for isomorphic graphs to uncover intrinsic interactions within biological networks. By creating data augmentations through vertex permutations, we train models to learn permutation-invariant representations. ResultsIn this study, we propose a contrastive learning approach tailored for isomorphic graphs to uncover intrinsic interactions within biological networks. By creating data augmentations through vertex permutations, we train models to learn permutation-invariant representations. Our approach was validated using five cancer-targeting biomarkers: ADGRF5, TP53, BRAF, KRAS, and GNAS. ConclusionWe discovered new connections between G-coupled receptors (GPR137B, GPR161, and GPR27) and key path-ways, interactions between cyclin-dependent kinase inhibitors (CDKN1A and CDK8) and specific biomarkers, and identified NFK-BIA as a central node linking all targeting biomarkers. This study highlights the potential of contrastive learning to reveal novel insights into cancer research and therapeutic targets. The implementation of this project is made available at: https://github.com/namnguyen0510/Contrastive-Learning-for-Graph-Based-Biological-Interaction-Discovery.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.