Back

BMDD: A Probabilistic Framework for Accurate Imputation of Zero-inflated Microbiome Sequencing Data

Zhou, H.; Chen, J.; Zhang, X.

2025-05-12 genomics
10.1101/2025.05.08.652808 bioRxiv
Show abstract

Microbiome sequencing data are inherently sparse and compositional, with excessive zeros arising from biological absence or insufficient sampling. These zeros pose significant challenges for downstream analyses, particularly those that require log-transformation. We introduce BMDD (BiModal Dirichlet Distribution), a novel probabilistic modeling framework for accurate imputation of microbiome sequencing data. Unlike existing imputation approaches that assume unimodal abundance, BMDD captures the bimodal abundance distribution of the taxa via a mixture of Dirichlet priors. It uses variational inference and a scalable expectation-maximization algorithm for efficient imputation. Through simulations and real microbiome datasets, we demonstrate that BMDD outperforms competing methods in reconstructing true abundances and improves the performance of differential abundance analysis. Through multiple posterior samples, BMDD enables robust inference by accounting for uncertainty in zero imputation. Our method offers a principled and computationally efficient solution for analyzing high-dimensional, zero-inflated microbiome sequencing data and is broadly applicable in microbial biomarker discovery and host-microbiome interaction studies. BMDD is available at: https://github.com/zhouhj1994/BMDD. Author SummaryUnderstanding the microbes living in and on our bodies--the microbiome--relies on analyzing complex sequencing data. However, these data often contain many zeros, either because a microbe is truly absent or simply missed due to insufficient sampling. These missing values make it hard to accurately analyze microbial patterns and identify important differences between groups, especially for methods that work on a log scale. To address this, we developed a new method called BMDD that uses a more realistic model to impute the zeros. Unlike existing tools that assume each microbe follows an unimodal abundance distribution, BMDD allows for microbes to follow a bimodal distribution, so they could behave differently in different conditions. It provides not just a single guess, but a range of possible values to better reflect the uncertainty. Our testing shows that BMDD more accurately recovers the true microbial profiles and improves the ability to detect meaningful differences between groups. This method can help researchers gain clearer insights into how the microbiome affects health and disease.

Published in PLOS Computational Biology (predicted rank #4) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Biometrics
23 papers in training set
Top 0.1%
26.8%
2
Bioinformatics
1204 papers in training set
Top 0.9%
22.6%
3
Molecular Ecology Resources
171 papers in training set
Top 0.3%
6.3%
50% of probability mass above
PLOS Computational Biology · published here
1863 papers in training set
Top 6%
6.3%
5
The Annals of Applied Statistics
19 papers in training set
Top 0.1%
3.5%
6
Biostatistics
24 papers in training set
Top 0.1%
3.5%
7
G3: Genes, Genomes, Genetics
252 papers in training set
Top 1%
3.3%
8
Methods in Ecology and Evolution
176 papers in training set
Top 0.8%
2.4%
9
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.1%
10
Genome Biology
637 papers in training set
Top 5%
1.7%
11
BMC Genomics
406 papers in training set
Top 4%
1.7%
12
BMC Bioinformatics
457 papers in training set
Top 4%
1.5%
13
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
14
GENETICS
483 papers in training set
Top 3%
1.1%
15
Genetic Epidemiology
55 papers in training set
Top 0.6%
1.0%
16
Nature Communications
5641 papers in training set
Top 54%
1.0%
17
Molecular Biology and Evolution
542 papers in training set
Top 5%
1.0%
18
Journal of Computational Biology
48 papers in training set
Top 1.0%
0.9%
19
Genome Research
468 papers in training set
Top 6%
0.9%
20
PeerJ
308 papers in training set
Top 11%
0.9%
21
G3 Genes|Genomes|Genetics
351 papers in training set
Top 4%
0.9%
22
Peer Community Journal
281 papers in training set
Top 5%
0.9%
23
Statistics in Medicine
40 papers in training set
Top 0.7%
0.6%