A multivariate method to correct for batch effects in microbiome data
Wang, Y.; Le Cao, K.-A.
Show abstract
Microbial communities are highly dynamic and sensitive to changes in the environment. Thus, microbiome data are highly susceptible to batch effects, defined as sources of unwanted variation that are not related to, and obscure any factors of interest. Existing batch correction methods have been primarily developed for gene expression data. As such, they do not consider the inherent characteristics of microbiome data, including zero inflation, overdispersion and correlation between variables. We introduce a new multivariate and non-parametric batch correction method based on Partial Least Squares Discriminant Analysis. PLSDA-batch first estimates treatment and batch variation with latent components to then subtract batch variation from the data. The resulting batch effect corrected data can then be input in any downstream statistical analysis. Two variants are also proposed to handle unbalanced batch x treatment designs and to include variable selection during component estimation. We compare our approaches with existing batch correction methods removeBatchEffect and ComBat on simulated and three case studies. We show that our three methods lead to competitive performance in removing batch variation while preserving treatment variation, and especially when batch effects have high variability. Reproducible code and vignettes are available on GitHub.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Ensuring that fundamentals of quantitative microbiology are reflected in microbial diversity analyses based on next-generation sequencing 93%
- ARGContextProfiler: Extracting and Scoring the Genomic Contexts of Antibiotic Resistance Genes using Assembly Graphs 92%
- Interpretations of microbial community studies are biased by the selected 16S rRNA gene amplicon sequencing pipeline. 92%
Similar papers in this journal
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 94%
- Comprehensive evaluation of methods for differential expression analysis of metatranscriptomics data 94%
- A powerful framework for an integrative study with heterogeneous omics data: from univariate statistics to multi-block analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.