Application of Attention and Graph Transformer-Based Approaches for RNA Biomarker Discovery in Metabolically-Associated Fatty Liver Disease (MAFL/NASH)
Cheruvu, A. S.; Zezulinski, D.; Sayeed, A.
Show abstract
The prevalence of nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH) in the United States has reached epidemic proportions, increasing the risk of liver cirrhosis and cancer. Current methods of diagnosis for NAFLD/NASH are invasive and costly, motivating the need for genetic "RNA" biomarkers detectable in a blood sample. In this study, explainable artificial intelligence (XAI) techniques are employed to increase the interpretability of the deep learning models in detecting the potential mRNA biomarker candidates for NAFLD/NASH. Nine RNA datasets ([~]1000 patients) with NAFLD/NASH were collected from the Gene Expression Omnibus. After conducting a differential gene expression analysis to reduce the dimensionality of the expression data, single-head and multi-head attention models were compared to baseline machine learning models in their ability to classify patients as NAFLD/NASH/healthy. XAI methods, including L1 regularization on baseline models and analysis of the internal attention matrix of the attention models, were utilized to identify biomarker candidates based on the relative importance of genes. The attention models achieved superior performance (accuracy: 67.5%) compared to the baseline models (Negative Binomial Linear Discriminant Analysis-62.64%; Poisson Linear Discriminant Analysis with Power Transformation - 58.24%). The top 17 and top 20 XAI-identified biomarkers with the baseline machine learning algorithms and the attention-based models respectively were then evaluated in lab. Preliminary data from in-lab validation confirmed upregulation of MT-ND3, HLA-B, APOC-1, and APOL-1 in NAFLD/NASH patients. Attention models have shown promise in identifying expression-based mRNA biomarkers and accurately diagnosing patients with NAFLD/NASH.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhanced Annotation of CD45RA to Distinguish T cell Subsets in Single Cell RNA-seq via Machine Learning 94%
- Identification of Monotonically Classifying Pairs of Genes for Ordinal Disease Outcomes 93%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 93%
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 94%
- Use of a graph neural network to the weighted gene co-expression network analysis of Korean native cattle 94%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 94%
Similar papers in this journal
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 95%
- Multi-omics Data Integration by Generative Adversarial Network 95%
- Integration of multi-source gene interaction networks and omics data with graph attention networks to identify novel disease genes 94%
Similar papers in this journal
- Single-Cell Classification Using Graph Convolutional Networks 94%
- De novo identification of maximally deregulated subnetworks based on multi-omics data with DeRegNet 94%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 94%
Similar papers in this journal
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 94%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 94%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.