Back

A Realistic Simulation Framework for EvaluatingMicrobiome Normalization in Sample Stratification and Differential Abundance

Al Khafaji, A. K.; Llorente, C. G.; Camacho, J.

2026-01-13 bioinformatics
10.64898/2026.01.13.699216 bioRxiv
Show abstract

BackgroundNormalization is a critical yet often poorly understood step in microbiome studies. Suboptimal approaches may lead to inaccurate conclusions in downstream analyses of microbial communities. Currently, there is no benchmarking framework to evaluate how normalisation affects both sample stratification and differential abundance simultaneously across taxonomic levels. In this paper, we propose a simulation pipeline based on real data and multivariate exploratory data analysis to provide a structured and reproducible assessment of normalization methods. ResultsNormalization methods exhibited distinct accuracy across taxonomic levels and sequencing depths. In our case study, at the phylum level, edgeR-TMM and Rarefaction improved accuracy by reducing coverage-related variation while preserving biological structure. In contrast, at the genus level, the overall improvement by normalization was less pronounced, reflecting the weaker influence of sequencing depth variability in this scenario, and EdgeR-TMM again provided the most accurate estimation of biological effect. Multivariate visualizations supported these observations, highlighting both sample-level and taxon-level differences among methods. Yet, ordination-based summaries are not sufficient for differential abundance inference and can be misleading, motivating the use of a simulation environment with known ground truth. ConclusionsNormalization performance varied with sequencing depth, sparsity, taxonomic resolution, and dataset size. Thus, there is no single normalization method that is expected to be optimal across all conditions. Our proposed simulation and analysis framework offers a reproducible and interpretable platform to evaluate existing and new normalization approaches in microbiome research for specific case studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.