Back

Feature matrix normalization, transformation and calculation of beta-diversity in metagenomics: Theoretical and applied perspectives on your decisions

Poulsen, C. S.; Aarestrup, F. M.; Brinch, C.; Ekstroem, C. T.

2019-11-29 microbiology
10.1101/859157 bioRxiv
Show abstract

Microbial metagenomics utilising next generation sequencing is a powerful experimental approach enabling detailed and potentially complete descriptions of the microbial world around and within us. Selecting how to perform feature data normalization, transformation and calculate {beta}-diversity is a critical step in the analysis of metagenomic data, but also a step for which a multitude of methods are available. Researchers need to have a broad overview and understand the many methods that exist in the field and the consequences from applying them. In this perspectives article, some of the most widely used metagenomic feature data normalizations, transformations and {beta}-diversity metrics are discussed in the context of multivariate visualizations. We provide a framework that other researchers can utilize to evaluate how robust their test data are when applying different normalizations, transformations and {beta}-diversity metrics, and visually compare the results of the methods. We constructed an in silico test dataset to evaluate the setup and clarify how the theoretical discussion is transferable to this data. We urge other researchers to implement their own test data, normalization, transformation, {beta}-diversity metric and visualization methods, in the hope that it will advance better decision making both in study design and analysis strategy.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.