Back

Deep-learning-based interpolation of longitudinal microbiome data powers biologically informative discovery

Qu, Y.; Lyu, R.; Wang, D.; Dai, Y.; Turcan, A.; Yu, S.; Xie, J.; Roach, J.; Butler, C.; Yap, P.-T.; Zhu, H.; Dashper, S.; Ribeiro, A. A.; Li, D.; Divaris, K.; Wu, D.

2025-02-17 microbiology
10.1101/2025.02.17.638709 bioRxiv
Show abstract

The human microbiome is a foundational and dynamic foundation for several health-related functions and disease processes. Advances in microbiome sequencing have enabled the characterization of microbial communities in several niches. Longitudinal microbiome studies further strive to discover clinically informative microbial community trajectories. However, these data are fraught with dropout events, high noise, and irregular sampling that limit and prevent the use of many available longitudinal analysis tools. To address these challenges, we introduce Bidirectional GRU-ODE-Bayes (BGOB), a deep learning framework developed for longitudinal microbiome interpolation. BGOB combines bidirectional information flow and ODE-based continuous modeling to jointly interpolate and smoothen trends across individual participants, providing uniform, denoised time intervals across patients. BGOB enables vastly improved performance in differential abundance testing and time-to-event analysis, and makes possible longitudinal analyses requiring uniformity, such as lead-lag detection and temporal clustering. After interpolation, previously low-powered datasets are able to broadly recapitulate known microbiology and elucidate interacting microbial communities. We highlight several associations between microbial taxa and disease, including novel species associated with Early Childhood Caries and disruption of key healthy gut microbiota in Inflammatory Bowel Disease. The BGOB package is publicly available at https://github.com/Rachel-Lyu/BGOB_n_test.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.