Back

Cross-tissue Graph Attention Networks for Semi-supervised Gene Expression Prediction

Wang, S.; He, M.; Qin, M.; Hu, Y.; Zhao, L.; Qin, Z.

2024-11-17 bioinformatics
10.1101/2024.11.15.623881 bioRxiv
Show abstract

High-throughput biotechnologies have significantly advanced precision medicine by enabling the exploitation of global gene expression patterns to enhance our understanding of disease etiology, progression, and treatment options. However, the tissue-specific nature of gene expression presents a challenge, particularly for less accessible tissues such as the brain, underscoring the need for computational methods to accurately impute gene expression in these critical but hard-to-reach tissues. While several attempts to impute gene expression in tissue-specific contexts have shown promising results, their reliance on regression analysis faces limitations due to the inability to capture complex, nonlinear relationships in gene expression patterns. In contrast, modern machine learning techniques, particularly graph neural networks, have demonstrated superior performance by efficiently modeling the intricate interactions among genes across different tissues. Therefore, we introduce gene expression imputation with Graph Attention Networks (gemGAT), a novel approach leveraging Graph Attention Networks (GATs) to enhance gene expression prediction across different tissues. gemGAT distinguishes itself by predicting the expression of all genes simultaneously, utilizing the full spectrum of genomic data to account for gene co-expressions and non-linear relationships. Validated through extensive experiments with Genotype-Tissue Expression (GTEx) data and a case study from the Alzheimers Disease Neuroimaging Initiative (ADNI), gemGAT demonstrates superior performance over existing methods by efficiently capturing non-linear gene co-expressions. This advancement underscores gemGATs potential to significantly contribute to precision medicine, showcasing its utility in advancing our understanding of gene expression in less accessible tissues.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.