Finding phylogeny-aware and biologically meaningful averages of metagenomic samples: L2UniFrac
Wei, W.; Millward, A. J.; Koslicki, D.
Show abstract
Metagenomic samples have high spatiotemporal variability. Hence, it is useful to summarize and characterize the microbial makeup of a given environment in a way that is biologically reasonable and interpretable. The UniFrac metric has been a robust and widely-used metric for measuring the variability between metagenomic samples. We propose that the characterization of metagenomic environments can be achieved by finding the average, a.k.a. the barycenter, among the samples with respect to the UniFrac distance. However, it is possible that such a UniFrac-average includes negative entries, making it no longer a valid representation of a metagenomic community. To overcome this intrinsic issue, we propose a special version of the UniFrac metric, termed L2UniFrac, which inherits the phylogenetic nature of the traditional UniFrac and with respect to which one can easily compute the average, producing biologically meaningful environment-specific "representative samples". We demonstrate the usefulness of such representative samples as well as the extended usage of L2UniFrac in efficient clustering of metagenomic samples, and provide mathematical characterizations and proofs to the desired properties of L2UniFrac. A prototype implementation is provided at: https://github.com/KoslickiLab/L2-UniFrac.git.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaMLP: A fast word embedding based classifier to profile target gene databases in metagenomic samples 95%
- Metabolic pathway prediction using non-negative matrix factorization with improved precision 94%
- Determining significant correlation between pairs of extant characters in a small parsimony framework 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Hierarchical non-negative matrix factorization using clinical information for microbial communities. 97%
- Machine learning based imputation techniques for estimating phylogenetic trees from incomplete distance matrices 96%
- Microbial trend analysis for common dynamic trend, group comparison and classification in longitudinal microbiome study 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.