RSim: A Reference-Based Normalization Method via Rank Similarity
Yuan, B.; Wang, S.
Show abstract
Microbiome sequencing data normalization is crucial for eliminating technical bias and ensuring accurate downstream analysis. However, this process can be challenging due to the high frequency of zero counts in microbiome data. We propose a novel reference-based normalization method called normalization via rank similarity (RSim) that corrects sample-specific biases, even in the presence of many zero counts. Unlike other normalization methods, RSim does not require additional assumptions or treatments for the high prevalence of zero counts. This makes it robust and minimizes potential bias resulting from procedures that address zero counts, such as pseudo-counts. Our numerical experiments demonstrate that RSim reduces false discoveries, improves detection power, and reveals true biological signals in downstream tasks such as PCoA plotting, association analysis, and differential abundance analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-scale Adaptive Differential Abundance Analysis in Microbial Compositional Data 98%
- Zero is not absence: censoring-based differential abundance analysis for microbiome data 97%
- Estimating and testing the microbial causal mediation effect with high-dimensional and compositional microbiome data 96%
Similar papers in this journal
Similar papers in this journal
- HiCzin: Normalizing metagenomic Hi-C data and detecting spurious contacts using zero-inflated negative binomial regression 94%
- MetaMLP: A fast word embedding based classifier to profile target gene databases in metagenomic samples 93%
- Metabolic pathway prediction using non-negative matrix factorization with improved precision 92%
Similar papers in this journal
- Dirichlet-multinomial modelling outperforms alternatives for analysis of microbiome and other ecological count data 95%
- On the impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters 93%
- Assessment of current taxonomic assignment strategies for metabarcoding eukaryotes 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.