Comprehensive analysis of network reconstruction approaches based on correlation in metagenomic data
Fuschi, A.; Merlotti, A.; Tran, D.-B.; Nguyen, H.; Weinstock, G.; Remondini, D.
Show abstract
Microbiome analysis is transforming our understanding of biological processes related to human health, epidemiology (antimicrobial resistance, horizontal gene transfer) environmental and agricultural studies. At the core of microbiome analysis is the description of microbial communities based on quantification of microbial taxa and dynamics. In the study of bacterial abundances, it is becoming more relevant to consider their relationship, to embed these data in the framework of network theory, allowing characterization of features like node relevance, pathway and community structure. In this work we characterize the principal biases in reconstructing networks from correlation measures, associated with the compositional character of relative abundance data, the diversity of abundances and the presence of unobserved species within a single sample, that might lead to wrong correlation estimates. We show how most of these problems can be overcome by applying typical transformations for compositional data, that allow the application of simple measures such as Pearsons correlation to correctly identify the positive and negative relationships between relative abundances, when data dimensionality is sufficiently high. Some issues remain, like the role of data sparsity, that if not properly addressed can lead to imbalances in correlation coefficient distribution.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NetCoMi: Network Construction and Comparison for Microbiome Data in R 95%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 95%
- Comprehensive evaluation of methods for differential expression analysis of metatranscriptomics data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.