Alpha and beta-diversities performance comparison between different normalization methods and centered log-ratio transformation in a microbiome public dataset
Bars-Cortina, D.
Show abstract
Microbiome data obtained after ribosomal RNA or shotgun sequencing represent a challenge for their ecological and statistical interpretation. Microbiome data is compositional data, with a very different sequencing depth between sequenced samples from the same experiment and harboring many zeros. To overcome this scenario, several normalizations and transformation methods have been developed to correct the microbiome datas technical biases, statistically analyze these data more optimally, and obtain more confident biological conclusions. Most existing studies have compared the performance of different normalization methods mainly linked to microbial differential abundance analysis methods but without addressing the initial statistical task in microbiome data analysis: alpha and beta-diversities. Furthermore, most of the studies used simulated microbiome data. The present study attempted to fill this gap. A public whole shotgun metagenomic sequencing dataset from a USA cohort related to gastrointestinal diseases has been used. Moreover, the performance comparison of eleven normalization methods and the transformation method based on the centered log ratio (CLR) has been addressed. Two strategies were followed to attempt to evaluate the aptitude of the normalization methods between them: the centered residuals obtained for each normalization method and their coefficient of variation. Concerning alpha diversity, the Shannon-Weaver index has been used to compare its output to the normalization methods. Regarding beta-diversity (multivariate analysis), it has been explored three types of analysis: principal coordinate analysis (PCoA) as an exploratory method; distance-based redundancy analysis (db-RDA) as interpretative analysis; and sparse Partial Least Squares Discriminant Analysis (sPLS-DA) as machine learning discriminatory multivariate method. Moreover, other microbiome statistical approaches were compared along the normalization and transformation methods: permutational multivariate analysis of variance (PERMANOVA), analysis of similarities (ANOSIM), beta-dispersion and multi-level pattern analysis in order to associate specific species to each type of diagnosis group in the dataset used. The GMPR (geometric mean of pairwise ratios) normalization method presented the best results regarding the dispersion of the new matrix obtained after being scaled. For the case of diversity, no differences were detected among the normalization methods compared. In terms of {beta} diversity, the db-RDA and the sPLS-DA analysis have allowed us to detect the most meaningful differences between the normalization methods. The CLR transformation method was the most informative in biological terms, allowing us to make more predictions. Nonetheless, it is important to emphasize that the CLR method and the UQ normalization method have been the only ones that have allowed us to make predictions from the sPLS-DA analysis, so their use could be more encouraged.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 95%
- Gene co-expression network analysis of the human gut commensal bacterium Faecalibacterium prausnitzii based on WGCNA in R-Shiny 95%
- Phylogeny analysis of whole protein-coding genes in metagenomic data detected an environmental gradient for the microbiota 94%
Similar papers in this journal
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 97%
- Meta-analysis of Microbiome Association Networks Reveal Patterns of Dysbiosis in Diseased Microbiomes 95%
- Manually weighted taxonomy classifiers improve species-specific rumen microbiome analysis compared to unweighted or average weighted taxonomy classifiers 95%
Similar papers in this journal
- Association of Body Index with Fecal Microbiome in Children Cohorts with Ethnic-Geographic Factor Interaction: Accurately Using a Bayesian Zero-inflated Negative Binomial Regression Model 97%
- parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis 96%
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 96%
Similar papers in this journal
Similar papers in this journal
- Hierarchical non-negative matrix factorization using clinical information for microbial communities. 97%
- Inferring directional relationships in microbial communities using signed Bayesian networks 96%
- Microbial trend analysis for common dynamic trend, group comparison and classification in longitudinal microbiome study 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.