Back

Expression Graph Network Framework For Biomarkerdiscovery

Liu, Y.; Kannan, K.; Huse, J. T.

2025-05-01 bioinformatics
10.1101/2025.04.28.651033 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWBiomarker discovery for complex diseases like cancer hinges on uncovering molecular signatures that capture intricate, interconnected relationships within biological data, a challenge that traditional statistical and machine learning methods often fail to meet due to the complexity of high-dimensional gene expression profiles. To overcome this, we introduce the Expression Graph Network Framework (EGNF), a cutting-edge graph-based approach that integrates graph neural networks (GNNs) with network-based feature engineering to enhance predictive biomarker identification. EGNF constructs biologically informed networks by combining gene expression data and clinical attributes within a graph database, utilizing hierarchical clustering to generate dynamic, patient-specific representations of molecular interactions. Leveraging graph learning techniques, including Graph Convolutional Networks (GCNs) and Graph Attention Networks (GATs), our framework identifies statistically significant and biologically relevant gene modules for classification. Validated across three independent datasets consisting of contrasting tumor types and clinical scenarios, EGNF consistently outperforms traditional machine learning models, achieving superior classification accuracy and interpretability. Notably, it delivers perfect separation between normal and tumor samples while excelling in nuanced tasks such as classifying disease progression and treatment outcome. This scalable, interpretable, and robust framework provides a powerful tool for biomarker discovery, with wide-ranging applications in precision medicine and the elucidation of disease mechanisms across diverse clinical contexts.

Published in Briefings in Bioinformatics (predicted rank #5) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.