waveome: a toolkit for longitudinal omics analysis using Gaussian processes
Ross, A.; Lloyd-Price, J.; Rahnavard, A.
Show abstract
Identifying meaningful associations from small-sample longitudinal data is challenging, especially in low signal-to-noise environments where the Gaussian likelihood assumption does not hold. We introduce two methods to algorithmically perform variable selection with sparse, irregularly sampled, longitudinal count data with over-dispersion to characterize nonlinear relationships between omics measurements and covariates of interest using Gaussian processes. The first is an additive non-greedy search-based method, while the second is a penalization approach using Horseshoe priors on kernel hyperparameters. In simulation studies, both methods outperform conventional statistical models in terms of distributional fit and exhibit a trade-off in feature selection. Applying the penalized variant to a real-world Crohns disease cohort, we recover well-established biomarkers, such as short-chain fatty acids, secondary bile acids, and specific lipid species, and uncover novel candidates for cross-sectional and temporal disease severity. Both methods are implemented in an open-source Python library, waveome, offering a robust set of tools for longitudinal biomarker discovery.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes 94%
- Semi-supervised Bayesian integration of multiple spatial proteomics datasets 94%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 94%
Similar papers in this journal
Similar papers in this journal
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 95%
- An integrated Bayesian framework for multi-omics prediction and classification 93%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.