Back

waveome: a toolkit for longitudinal omics analysis using Gaussian processes

Ross, A.; Lloyd-Price, J.; Rahnavard, A.

2026-03-05 bioinformatics
10.64898/2026.03.03.709362 bioRxiv
Show abstract

Identifying meaningful associations from small-sample longitudinal data is challenging, especially in low signal-to-noise environments where the Gaussian likelihood assumption does not hold. We introduce two methods to algorithmically perform variable selection with sparse, irregularly sampled, longitudinal count data with over-dispersion to characterize nonlinear relationships between omics measurements and covariates of interest using Gaussian processes. The first is an additive non-greedy search-based method, while the second is a penalization approach using Horseshoe priors on kernel hyperparameters. In simulation studies, both methods outperform conventional statistical models in terms of distributional fit and exhibit a trade-off in feature selection. Applying the penalized variant to a real-world Crohns disease cohort, we recover well-established biomarkers, such as short-chain fatty acids, secondary bile acids, and specific lipid species, and uncover novel candidates for cross-sectional and temporal disease severity. Both methods are implemented in an open-source Python library, waveome, offering a robust set of tools for longitudinal biomarker discovery.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.