CAFT: A Compositional Log-Linear Model for Microbiome Data with Zero Cells
Satten, G. A.; Li, M.; Zhao, N.
Show abstract
BackgroundDifferential abundance analysis is fundamental to microbiome research and provides valuable insights into host-microbe interactions. However, microbiome data are compositional, highly sparse (with many zero counts), and influenced by differential experimental biases across taxa. Standard statistical methods often overlook these features. Many approaches analyze relative abundances without accounting for compositionality or rely on pseudocounts, potentially leading to spurious associations and inadequate false discovery rate (FDR) control. MethodsWe introduce a novel framework for differential abundance analysis of microbiome data: the Compositional Accelerated Failure Time (CAFT) model. CAFT addresses zero read counts by treating them as censored observations that are below a detection limit. This approach is inherently resistant to multiplicative technical bias, eliminates the need for pseudocounts, and addresses compositional bias through the establishment of appropriate score test procedures. ResultsExtensive simulations show that CAFT outperforms competing compositional differential abundance methods, including LOCOM, LinDA, ANCOM-BC2, its robust variant, and LDM-clr by offering more robust type I error and FDR control with or without technical bias. Additionally, we applied CAFT to microbiome data on inflammatory bowel disease (IBD) and the upper respiratory tract (URT) to identify differentially abundant gut microbial taxa between IBD patients and healthy controls, as well as URT taxa distinguishing smokers from non-smokers. ConclusionWe present CAFT, a powerful, robust, and efficient approach for compositional differential abundance analysis. CAFT effectively controls Type I error and maintains FDR control, while demonstrating enhanced power in statistical testing. These capabilities render CAFT a useful tool for compositional microbiome data analysis. Availability and implementationThe R package and Vignette are available at https://github.com/mli171/CAFT.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Defining and Evaluating Microbial Contributions to Metabolite Variation in Microbiome-Metabolome Association Studies 94%
- A mixed model approach for estimating drivers of microbiota community composition and differential taxonomic abundance 94%
- parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis 94%
Similar papers in this journal
- Decoding the Language of Microbiomes: Leveraging Patterns in 16S Public Data using Word-Embedding Techniques and Applications in Inflammatory Bowel Disease 96%
- CBEA: Competitive balances for taxonomic enrichmentanalysis 96%
- Modeling the temporal dynamics of the gut microbial community in adults and infants 95%
Similar papers in this journal
- A multi-view model for relative and absolute microbial abundances 95%
- Compositional knockoff filter for high-dimensional regression analysis of microbiome data 95%
- A Mixed Effect Similarity Matrix Regression Model (SMRmix) for Integrating Multiple Microbiome Datasets at Community Level and its Application in HIV 94%
Similar papers in this journal
- Feature selection and causal analysis for microbiome studies in the presence of confounding using standardization 96%
- A negative binomial latent factor model for paired microbiome sequencing data 94%
- Functional Analysis of Metagenomes by Likelihood Inference (FAMLI) Successfully Compensates for Multi-Mapping Short Reads from Metagenomic Samples 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.