Back

Differential Expression Analysis for Metatranscriptomics with Sample-Paired Metagenomics Data

Yan, B.; Wu, D.

2022-12-11 genomics
10.1101/2022.12.08.519567 bioRxiv
Show abstract

1MotivationLarge microbial communities have contributed to essential ecosystems and multiple health and disease processes. Metatranscriptomics (MTX) is becoming increasingly important for profiling the gene expression pattern and functional activities of microbial communities. A fundamental task for analyzing the high-throughput sequencing data is the Differential Expression (DE) analysis, which aims to identify up/down-regulated genes/pathways across multiple conditions. However, DE analysis is complicated by the variation of underlying DNA copies, which is caused by the highly dynamic nature of microbial communities. The non-independent change between MTX and MGX may lead to false discoveries about which genes are biologically differentially expressed. Nevertheless, We can settle this problem by using the information from paired metagenomics (MGX) data. ResultsWe proposed the MetaDePair, a statistical model for DE analysis in MTX with paired MGX data. It relied on a conditional Negative-Binomial distribution to model the count data. We showed that the adjustment of underlying DNA copies could significantly eliminate the variations of RNA data, thus improving the power of statistical inference. Also, we found that appropriate pre-filtering of zeros can also improve the sensitivity and precision of the model. Moreover, we proposed a method to simulate the paired MTX and MGX data. Using simulated data and real microbial data sets, we demonstrated that our method has a high statistical power while controlling the false discovery rate (FDR). We applied the tool to oral microbiome data and identified significant genes that are associated with Early Childhood Caries (ECC). Our tool enabled a more accurate and effective DE analysis based on MTX with paired MGX data and helped us improve the understanding of the functional characterization of microbial communities.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.