Back

Understanding the bias of compositional microbiome differential abundance estimation

Calle, M. L.; Pujolassos, M.; Susin, A.

2026-04-30 bioinformatics
10.64898/2026.04.28.721392 bioRxiv
Show abstract

One of the most relevant objectives in microbiome studies is the identification of microbial species that are differentially abundant across conditions. However, the compositional nature of microbiome data complicates this task. Interdependence among components leads to spurious associations when the abundances of each component are analyzed separately. Due to the growing awareness of the challenges of compositional data analysis (CoDA), log-ratio transformations, such as the additive log-ratio (alr) or the centered log-ratio (clr) transformations, have become increasingly popular in microbiome studies. Several studies have compared the performance of compositional and non-compositional methods through simulations. However, the debate between these two frameworks remains unresolved, creating confusion among researchers. Rather than relying on simulation-based results, this work provides theoretical results that enable a more rigorous and conclusive analysis of the problem, contributing to a better understanding of differential abundance estimation. We provide theoretical expressions of the bias of differential abundance estimation related to the use of proportions (total sum scaling) and log-ratio transformations (alr and clr) when estimates are interpreted as absolute rather than relative to a reference. The factors that most strongly influence the bias are the magnitude and direction of the effects, the dimension of the composition, the proportion of differentially abundant variables, and the distribution of relative abundances. The findings of this work strongly support the use of CoDA transformations; however, they also highlight that even when log-ratio transformations are applied, interpreting the results outside of a CoDA framework can still lead to biased conclusions. Among CoDA transformations, alr has several advantages over clr: its reference is more explicit, which reduces the risk of interpreting estimates as absolute rather than relative, and it facilitates the replication of results in independent studies, as it only requires assessing changes relative to the same reference rather than reconstructing the full composition. In this work, we propose a heuristic method for selecting a suitable alr reference component, which will enable a more widespread use of this transformation.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 2%
12.8%
2
mSystems
394 papers in training set
Top 0.6%
9.8%
3
Bioinformatics
1204 papers in training set
Top 3%
8.0%
4
PLOS ONE
5266 papers in training set
Top 26%
6.3%
5
Journal of Computational Biology
48 papers in training set
Top 0.1%
5.6%
6
PeerJ
308 papers in training set
Top 0.7%
5.6%
7
BMC Bioinformatics
457 papers in training set
Top 2%
5.5%
50% of probability mass above
8
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.9%
9
Scientific Reports
3612 papers in training set
Top 26%
4.1%
10
BMC Genomics
406 papers in training set
Top 2%
3.3%
11
Microbiome
154 papers in training set
Top 1%
2.1%
12
Frontiers in Microbiology
427 papers in training set
Top 4%
2.1%
13
Biometrics
23 papers in training set
Top 0.1%
2.1%
14
Ecology and Evolution
267 papers in training set
Top 4%
1.7%
15
BMC Microbiology
49 papers in training set
Top 0.8%
1.5%
16
GigaScience
212 papers in training set
Top 3%
1.5%
17
Statistics in Medicine
40 papers in training set
Top 0.3%
1.5%
18
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
19
BioData Mining
22 papers in training set
Top 0.5%
1.1%
20
Nature Communications
5641 papers in training set
Top 56%
0.9%
21
Genes
144 papers in training set
Top 4%
0.9%
22
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.9%
23
Peer Community Journal
281 papers in training set
Top 5%
0.9%
24
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.6%
25
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%
26
FEMS Microbiology Ecology
54 papers in training set
Top 1%
0.6%
27
Physical Review E
112 papers in training set
Top 1%
0.6%
28
Journal of Theoretical Biology
162 papers in training set
Top 2%
0.6%