Integration of Proxy Intermediate Omics traits into a Nonlinear Two-Step model for accurate phenotypic prediction
Yoshioka, H.; Mary-Huard, T.; Aubert, J.; Toda, Y.; Ohmori, Y.; Yamasaki, Y.; Tsujimoto, H.; Takahashi, H.; Nakazono, M.; Takanashi, H.; Fujiwara, T.; Tsuda, M.; Kaga, A.; Inaba, J.; Fuji, Y.; Hirai, M. Y.; Nose, Y.; Kumaishi, K.; Usui, E.; Kobori, S.; Sato, T.; Narukawa, M.; Ichihashi, Y.; Iwata, H.
Show abstract
Intermediate omics traits, which mediate the effects of genetic variation on phenotypic traits, are increasingly recognised as valuable components of genetic evaluation. In particular, rhizosphere microbiota play a crucial role in plant health and productivity; however, their complex interactions with host genetics remain challenging to model. Although two-step modeling frameworks have been proposed to integrate intermediate omics traits into phenotype prediction, existing approaches do not incorporate nonlinear relationships between different omics layers. To address this, we have proposed a two-step phenotype prediction framework that integrates genomic, rhizosphere microbiome, and metabolome (meta-metabolome) data, while explicitly capturing omicsomics nonlinearities. The first step is to predict meta-metabolome traits from genetic and microbial features, thus effectively isolating them from the environmental noise. In this process, intermediate "proxy" omics traits are generated as general biological information to provide robust models. The second step utilises this "proxy" to enhance the accuracy of the phenotype prediction. We compared the linear model (Best Linear Unbiased Prediction, BLUP) and the nonlinear model (Random Forest, RF) at each step, as demonstrated through simulations and empirical analysis of a multi-omics soybean dataset in which nonlinear modeling captures intricate omics interactions. Notably, our approach enables phenotype prediction without requiring the original meta-metabolome data used in model training, thereby reducing reliance on costly omics measurements. This framework integrates intermediate omics traits into genomic prediction to improve prediction accuracy and provide solutions for deeper insights into plant-microbiome interactions.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 94%
- A pipeline for the reconstruction and evaluation of context-specific human metabolic models at a large-scale 94%
- MiMeNet: Exploring Microbiome-Metabolome Relationships using Neural Networks 93%
Similar papers in this journal
Similar papers in this journal
- Single sample pathway analysis in metabolomics: performance evaluation and application 94%
- SCOUR: A stepwise machine learning framework for predicting metabolite-dependent regulatory interactions 94%
- Feature selection and causal analysis for microbiome studies in the presence of confounding using standardization 94%
Similar papers in this journal
- Comprehensive evaluation of methods for differential expression analysis of metatranscriptomics data 94%
- A powerful framework for an integrative study with heterogeneous omics data: from univariate statistics to multi-block analysis 93%
- VBayesMM: Variational Bayesian neural network to prioritize important relationships of high-dimensional microbiome multiomics data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.