Explainable autoencoder-based representation learning for gene expression data
Yu, Y.; Kossinna, P.; Liao, W.; Zhang, Q.
Show abstract
Modern machine learning methods have been extensively utilized in gene expression data analysis. In particular, autoencoders (AE) have been employed in processing noisy and heterogenous RNA-Seq data. However, AEs usually lead to "black-box" hidden variables difficult to interpret, hindering downstream experimental validation and clinical translation. To bridge the gap between complicated models and biological interpretations, we developed a tool, XAE4Exp (eXplainable AutoEncoder for Expression data), which integrates AE and SHapley Additive exPlanations (SHAP), a flagship technique in the field of eXplainable AI (XAI). It quantitatively evaluates the contributions of each gene to the hidden structure learned by an AE, substantially improving the expandability of AE outcomes. By applying XAE4Exp to The Cancer Genome Atlas (TCGA) breast cancer gene expression data, we identified genes that are not differentially expressed, and pathways in various cancer-related classes. This tool will enable researchers and practitioners to analyze high-dimensional expression data intuitively, paving the way towards broader uses of deep learning. AvailabilityOpen source at https://github.com/QingrunZhangLab/Explainable-Deep-Autoencoder. Contactsqingrun.zhang@ucalgary.ca and wliao@ucalgary.ca.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Style transfer with variational autoencoders is a promising approach to RNA-Seq data harmonization and analysis 96%
- Adversarial Deconfounding Autoencoder for Learning Robust Gene Expression Embeddings 95%
- ECMarker: Interpretable machine learning model identifies gene expression biomarkers predicting clinical outcomes and reveals molecular mechanisms of human disease in early stages 95%
Similar papers in this journal
- Deepprune: Learning efficient and interpretable convolutional networks through weight pruning for predicting DNA-protein binding 93%
- Machine Learning Approaches Identify Genes Containing Spatial Information from Single-Cell Transcriptomics Data. 93%
- A Network-centric Framework for theEvaluation of Mutual Exclusivity Tests onCancer Drivers 93%
Similar papers in this journal
- Beyond synthetic lethality in large-scale metabolic and regulatory network models via genetic minimal intervention sets 93%
- Mining hidden knowledge: Embedding models of cause-effect relationships curated from the biomedical literature 93%
- SPREd: A simulation-supervised neural network tool for gene regulatory network reconstruction 93%
Similar papers in this journal
- Accurate Prediction of Virus-Host Protein-Protein Interactions via a Siamese Neural Network Using Deep Protein Sequence Embeddings 93%
- MUSTANG: MUlti-sample Spatial Transcriptomics data ANalysis with cross-sample transcriptional similarity Guidance 93%
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.