Back

Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding

Lee, E. M.; McNulty, N. P.; Cheng, J.; Chang, H.-W.; Hibberd, M. C.; Cohen, B. A.; Gordon, J.

2025-09-16 genomics
10.1101/2025.09.11.675436 bioRxiv
Show abstract

Metatranscriptomic (MTX) sequencing quantifies gene expression from the collective genomes of microbial communities (microbiomes), enabling assessment of functional activity rather than functional potential. While differential expression testing is essential for RNA-sequencing analysis, current metatranscriptomic approaches have only been benchmarked on simulated data, resulting in a lack of standard practices for analysis of real datasets. Here, we use mock communities (defined mixtures of microbial cells with known properties) to quantitatively assess robustness and susceptibility of current approaches to various confounders including organisms low relative abundance, differential abundance, low prevalence, global transcriptional output changes, and compositional effects. We show that no current method is robust to all confounders and method performance on simulated data does not generalize to real datasets. We then apply the same approaches to MTX datasets generated from gnotobiotic mice colonized with defined consortia of human bacterial strains and show that the method nominated by the mock community comparisons successfully inferred cross-feeding dynamics that were subsequently validated in vitro. Finally, using metagenome-assembled genomes from a human clinical study, we leverage genome-level sequencing depth and detection of genes to exclude low information samples on a per-organism basis to overcome confounding low prevalence and enhance differential expression inference.

Published in Nature Communications (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.