Trait genetic architecture and population structure determine model selection for genomic prediction in natural Arabidopsis Thaliana populations.
Gibbs, P.; Paril, J.; Fournier-Level, A.
Show abstract
Genomic prediction applies to a wide range of agronomically relevant traits, with distinct ontologies and genetic architectures. Selecting the most appropriate model for the distribution of genetic effects and their associated allele frequencies in the training population is crucial. Linear regression models are often preferred for genomic prediction. However, linear models may not suit all genetic architectures and training populations. Machine Learning approaches have been proposed to improve genomic prediction owing to their capacity to capture complex biology including epistasis. However, the applicability of different genomic prediction models, including non-linear/non-parametric approaches, have not been rigorously assessed across a wide variety of plant traits in natural outbreeding populations. This study evaluates genomic prediction sensitivity to trait ontology and the impact of population structure on model selection and prediction accuracy. Examining 36 quantitative traits measured for 1000+ natural genotypes of the model plant Arabidopsis thaliana, we assessed the performance of penalised regression, random forest, and multilayer perceptron at producing genomic predictions. Regression models were generally the most accurate, except for biochemical traits where random forest performed best. We link this result to the genetic architecture of each trait - notably that biochemical traits have simpler genetic architecture than macroscopic traits. Moreover, complex macroscopic traits, particularly those related to flowering and yield, were strongly correlated to population structure, while molecular traits were better predicted by fewer, independent markers. This study highlights the relevance of machine learning approaches for simple molecular traits and underscores the need to consider ancestral population history when designing training samples. Article summaryMachine learning and linear models were tested for genomic prediction of multiple traits in the model plant Arabidopsis thaliana. We associate the performance of genomic prediction models to trait ontology, finding machine learning approaches applicable to biochemical traits, and linear models best for macroscopic traits. We link this result to the genetic architecture of each trait and patterns of selection in the association panels ancestral population, thus underscoring the relevance of these two sensitivities to genomic prediction in plant breeding.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Low additive genetic variation in a trait under selection in domesticated rice 96%
- Analysis of genotype by environment interactions in a maize mapping population 96%
- Genome-wide association and prediction study in grapevine deciphers the genetic architecture of multiple traits and identifies genes under many new QTLs 95%
Similar papers in this journal
Similar papers in this journal
- Genomic mating in outbred species: predicting cross usefulness with additive and total genetic covariance matrices 95%
- Pervasive GxE interactions shape adaptive trajectories and the exploration of the phenotypic space in artificial selection experiments 95%
- MegaLMM improves genomic predictions in new environments using environmental covariates 95%
Similar papers in this journal
- Neighbor GWAS: incorporating neighbor genotypic identity into genome-wide association studies of field herbivory 95%
- Physical geography, isolation by distance and environmental variables shape genomic variation of wild barley (Hordeum vulgare L. ssp. spontaneum in the Southern Levant 95%
- Revisiting a GWAS peak in Arabidopsis thaliana reveals possible confounding by genetic heterogeneity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.