Back

An Interpretable Sparse Graph Contrastive Learning Approach for Identifying Breast Cancer Risk Variants

Gudhe, N. R.; Hartikainen, J. M.; Tengstrom, M.; Pylkas, K.; Winqvist, R.; Kosma, V.-M.; Behravan, H.; Mannermaa, A.

2025-01-15 genetic and genomic medicine
10.1101/2025.01.13.25320451 medRxiv
Show abstract

Genome-wide association studies (GWASs) have identified over 2,400 genetic variants associated to breast cancer. Conventional GWASs methods that analyze variants independently often overlook the complex genetic interactions underlying disease susceptibility. Machine and deep learning approaches present promising alternatives, yet encounter challenges, including overfitting due to high dimensionality ([~]10 million variants) and limited sample sizes, as well as limited interpretability. Here, we present GenoGraph, a graph-based contrastive learning framework designed to address these limitations by modeling high-dimensional genetic data in low-sample-size scenarios. We demonstrate GenoGraphs efficacy in breast cancer case-control classification task, achieving accuracy of 0.96 using the Biobank of Eastern Finland dataset. GenoGraph identified rs11672773 (ZNF8) as a key risk variant in Finnish population, with significant interactions with rs10759243 (KLF4) and rs3803662 (TOX3). Furthermore, in silico validation confirmed the biological relevance of these findings, underscoring GenoGraphs potential to advance breast cancer risk prediction and elucidate genetic interactions for personalized medicine.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.