A semi-parametric multiple imputation method for high-sparse, high-dimensional, compositional data
Sohn, M. B.; Scheible, K.; Gill, S. R.
Show abstract
High sparsity (i.e., excessive zeros) in microbiome data, which are high-dimensional and compositional, is unavoidable and can significantly alter analysis results. However, efforts to address this high sparsity have been very limited because, in part, it is impossible to justify the validity of any such methods, as zeros in microbiome data arise from multiple sources (e.g., true absence, stochastic nature of sampling). The most common approach is to treat all zeros as structural zeros (i.e., true absence) or rounded zeros (i.e., undetected due to detection limit). However, this approach can underestimate the mean abundance while overestimating its variance because many zeros can arise from the stochastic nature of sampling and/or functional redundancy (i.e., different microbes can perform the same functions), thus losing power. In this manuscript, we argue that treating all zeros as missing values would not significantly alter analysis results if the proportion of structural zeros is similar for all taxa, and we propose a semi-parametric multiple imputation method for high-sparse, high-dimensional, compositional data. We demonstrate the merits of the proposed method and its beneficial effects on downstream analyses in extensive simulation studies. We reanalyzed a type II diabetes (T2D) dataset to determine differentially abundant species between T2D patients and non-diabetic controls.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero is not absence: censoring-based differential abundance analysis for microbiome data 97%
- Estimating and testing the microbial causal mediation effect with high-dimensional and compositional microbiome data 97%
- A Rarefaction-Based Extension of the LDM for Testing Presence-Absence Associations in the Microbiome 96%
Similar papers in this journal
- Dirichlet-multinomial modelling outperforms alternatives for analysis of microbiome and other ecological count data 96%
- On the impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters 94%
- Assessment of current taxonomic assignment strategies for metabarcoding eukaryotes 93%
Similar papers in this journal
- A multi-view model for relative and absolute microbial abundances 96%
- Compositional knockoff filter for high-dimensional regression analysis of microbiome data 95%
- A Mixed Effect Similarity Matrix Regression Model (SMRmix) for Integrating Multiple Microbiome Datasets at Community Level and its Application in HIV 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.