Back

Predicting complex phenotypes using multi-omics data in maize

Creach, M.; Webster, B.; Newton, L.; Turkus, J.; Schnable, J.; Thompson, A.; VanBuren, R.

2025-10-01 plant biology
10.1101/2025.09.30.679283 bioRxiv
Show abstract

Understanding and predicting complex traits in plants remains a fundamental challenge due to the emergent nature of most phenotypes and their dependence on genetic, regulatory, and environmental interactions. Accurate prediction of traits and identification of underlying genetic elements has broad applications for plant breeding, systems biology, and biotechnology. Here, we tested if multi-omic datasets could improve predictive accuracy of 129 diverse maize phenotypes across nine environments using genomic markers, field based transcriptomic data from two locations, and drone-derived phenomic data of vegetative indices. We trained and compared linear (rrBLUP) and nonlinear (support vector regression) models using single- and multi-omics inputs. Multi-omics models consistently outperformed single-omics models for most traits, with genomic and transcriptomic inputs contributing distinct biological features. Phenomic features alone yielded the lowest predictive power but improved predictions for specific trait categories like root architecture. Transcriptomic datasets enabled cross-environment prediction, demonstrating that gene expression patterns from one field site could accurately predict traits measured in another. Environment-specific expression of benchmark flowering time genes highlighted the value of transcriptomics in capturing genotype-by-environment (GxE) interactions not detectable through genomic data alone. These findings demonstrate that integrating transcriptomic and phenomic data with genotypes enhances trait prediction, improves model generalizability across environments, and provides deeper insight into the genetic and regulatory architecture of agriculturally important traits in maize.

Published in The Plant Cell (predicted rank #2) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.