A Computational Approach to Interpreting the Embedding Space of Dimension Reduction
Zhang, B.; Uno, K.; Kodama, H.; Himori, K.; Matsui, Y.
Show abstract
Nonlinear dimension reduction methods are widely applied in studies analyzing gene and protein expression, by revealing patterns of discrete groups and continuous orders in high-dimensional data. However, the tools are limited to understanding the obtained embedding structures of biological mechanisms, hindering the full exploitation of data. Here, we propose a novel framework to interpret embedding systematically by identifying and mapping associated biological functions. The method performs statistical tests and visualizes significantly enriched functions essential for the organization of the embedding structure, by applying it to the embedding results of two datasets: the Genotype Tissue Expression dataset and a Caenorhabditis elegans embryogenesis dataset, one capturing distinct cluster structures and the other capturing continuous developmental trajectories. We identified the associated functions for interpreting the two embeddings and confirmed it as a useful explainable AI tool in exploratory data analysis by providing annotations to the embedding space.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CENTRA: Knowledge-Based Gene Contexuality Graphs Reveal Functional Master Regulators by Centrality and Fractality 95%
- SIMPLEs: a single-cell RNA sequencing imputation strategy preserving gene modules and cell clusters variation 95%
- Transfer Learning Compensates Limited Data, Batch-Effects, And Technical Heterogeneity In Single-Cell Sequencing 95%
Similar papers in this journal
- GeneSurfer Enables Transcriptome-wide Exploration and Functional Annotation of Gene Co-expression Modules in 3D Spatial Transcriptomics Data 94%
- DeepBend: An Interpretable Model of DNA Bendability 94%
- BATMAN: fast and accurate integration of single-cell RNA-Seq datasets via minimum-weight matching 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.