A realistic benchmark for the identification of differentially abundant taxa in (confounded) human microbiome studies
Wirbel, J.; Essex, M.; Forslund, S. K.; Zeller, G.
Show abstract
BackgroundIn microbiome disease association studies, it is a fundamental task to test which microbes differ in their abundance between groups. Yet, consensus on suitable or optimal statistical methods for differential abundance (DA) testing is lacking, and it remains unexplored how these cope with confounding. Previous DA benchmarks relying on simulated datasets did not quantitatively evaluate the similarity to real data, which undermines their recommendations. ResultsHere we develop a simulation framework which implants calibrated signals into real taxonomic profiles, including signals mimicking confounders. Using several whole-metagenome and 16S rRNA gene amplicon datasets, we validate that our simulated data resembles real data from disease association studies to a much greater extent than in previous benchmarks. With extensively parametrized simulations we benchmark the performance of eighteen DA methods and further evaluate the best ones on confounded simulations. Only linear models, limma, fastANCOM, and the Wilcoxon test properly control false discoveries at relatively high sensitivity. When additionally considering confounders, these issues are exacerbated, but we find that post hoc adjustment can effectively mitigate them. In a large cardiometabolic disease dataset, we showcase that failure to account for covariates such as medication causes spurious association in real-world applications. ConclusionsFor microbiome association studies tight error control is critical. The unsatisfactory performance of many DA methods and the persistent danger of unchecked confounding suggest these contribute to a lack of reproducibility among such studies. We have open-sourced our simulation and benchmarking software to foster a much-needed consolidation of statistical methodology for microbiome research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of statistical methods from single cell, bulk RNA-seq and metagenomics applied to microbiome data 96%
- Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine-learning toolbox 95%
- Metalign: Efficient alignment-based metagenomic profiling via containment min hash 95%
Similar papers in this journal
Similar papers in this journal
- MCSPACE: inferring microbiome spatiotemporal dynamics from high-throughput co-localization data 96%
- Identifying unmeasured heterogeneity in microbiome data via quantile thresholding (QuanT) 96%
- Revealing Interactions between Microbes, Metabolites, and Dietary Compounds using Genome-scale Analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.