The GMD-biplot and its application to microbiome data
Wang, Y.; Randolph, T.; Shojaie, A.; Ma, J.
Show abstract
Exploratory analysis of human microbiome data is often based on dimension-reduced graphical displays derived from similarities based on non-Euclidean distances, such as UniFrac or Bray-Curtis. However, a display of this type, often referred to as the principal coordinate analysis (PCoA) plot, does not reveal which taxa are related to the observed clustering because the configuration of samples is not based on a coordinate system in which both the samples and variables can be represented. The reason is that the PCoA plot is based on the eigen-decomposition of a similarity matrix and not the singular value decomposition (SVD) of the sample-by-abundance matrix. We propose a novel biplot that is based on an extension of the SVD, called the generalized matrix decomposition (GMD), which involves an arbitrary matrix of similarities and the original matrix of variable measures, such as taxon abundances. As in a traditional biplot, points represent the samples and arrows represent the variables. The proposed GMD-biplot is illustrated by analyzing multiple real and simulated data sets which demonstrate that the GMD-biplot provides improved clustering capability and a more meaningful relationship between the arrows and the points.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Incorporate D-trace Loss into Compositional Data Analysis 96%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 95%
- Evaluating the Number of Different Genomes in a Metagenome by Means of the Compositional Spectra Approach 95%
Similar papers in this journal
- Extended Graphical Lasso for Multiple Interaction Networks for High Dimensional Omics Data 96%
- Over-optimism in unsupervised microbiome analysis: Insights from network learning and clustering 94%
- RCFGL: Rapid Condition adaptive Fused Graphical Lasso and application to modeling brain region co-expression networks 94%
Similar papers in this journal
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 96%
- Fast and robust imputation for miRNA expression data using constrained least squares 94%
- BayesianSSA: a Bayesian statistical model based on structural sensitivity analysis for predicting responses to enzyme perturbations in metabolic networks 93%
Similar papers in this journal
- Exploring High-Dimensional Biological Data with Sparse Contrastive Principal Component Analysis 95%
- Zero is not absence: censoring-based differential abundance analysis for microbiome data 94%
- Super-delta2: An Enhanced Differential Expression Analysis Procedure for Multi-Group Comparisons of RNA-seq Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.