Using Shared Features Improves Metabolite Effect Estimation
Dubey, H. V.; Farage, G.; Sen, S.
Show abstract
External biological knowledge provides valuable information about relationships among metabolites, yet this information is usually not incorporated directly into statistical estimation procedures. Most existing approaches estimate metabolite effects independently, ignoring known biochemical structure such as shared subclasses and pathway membership. We propose a Bayesian hierarchical framework that improves metabolite effect estimates by incorporating external biological information describing relationships among metabolites. The proposed method improves metabolite-specific estimates by allowing related metabolites to borrow information from one another while preserving metabolite-level inference. We evaluate the methodology using simulation studies across a range of sample sizes and heterogeneity regimes together with three metabolomics applications involving distinct biological annotation structures. Across both simulated and real datasets, incorporating external biological information consistently improves metabolite effect estimation. Gains are most pronounced when sample sizes are small and metabolite classes are informative, i.e. more homogenous within classes.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Single sample pathway analysis in metabolomics: performance evaluation and application 94%
- SCOUR: A stepwise machine learning framework for predicting metabolite-dependent regulatory interactions 93%
- AMON: Annotation of metabolite origins via networks to better integrate microbiome and metabolome data 93%
Similar papers in this journal
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 93%
- PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes 93%
- Pathway analysis in metabolomics: pitfalls and best practice for the use of over-representation analysis 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.