Back

Implications of taxonomic bias for microbial differential-abundance analysis

McLaren, M. R.; Nearing, J. T.; Willis, A. D.; Lloyd, K. G.; Callahan, B. J.

2022-08-19 bioinformatics
10.1101/2022.08.19.504330 bioRxiv
Show abstract

Differential-abundance (DA) analyses enable microbiome researchers to assess how microbial species vary in relative or absolute abundance with specific host or environmental conditions, such as health status or pH. These analyses typically use sequencing-based community measurements that are taxonomically biased to measure some species more efficiently than others. Understanding the effects that taxonomic bias has on the results of a DA analysis is essential for achieving reliable and translatable findings; yet currently, these effects are unknown. Here, we characterized these effects for DA analyses of both relative and absolute abundances, using a combination of mathematical theory and data analysis of real and simulated case studies. We found that, for analyses based on species proportions, taxonomic bias can cause significant errors in DA results if the average measurement efficiency of the community is associated with the condition of interest. These errors can be avoided by using more robust DA methods (based on species ratios) or quantified and corrected using appropriate controls. Wide adoption of our recommendations can improve the reproducibility, interpretability, and translatability of microbiome DA studies. This manuscript was rendered from commit 7412a36 of https://github.com/mikemc/differential-abundance-theory. Supporting data analyses can be found in the accompanying computational research notebook. Please post comments or questions on GitHub. The manuscript is licensed under a CC BY 4.0 License. See the GitHub Releases or Zenodo record for earlier versions.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.