Using visual scores and categorical data for genomic prediction of complex traits in breeding programs
Ferreira Azevedo, C.; Ferrao, L. F. V.; Benevenuto, J.; Resende, M. D.; Nascimento, M.; Nascimento, A. C.; Munoz, P.
Show abstract
Most genomic prediction methods are based on assumptions of normality due to their simplicity, robustness, and ease of implementation. However, in plant and animal breeding, target traits are often collected as categorical data, thus violating the normality assumption, which could affect the prediction of breeding values and the estimation of crucial genetic parameters. In this study, we examined the main challenges of categorical phenotypes in genomic prediction and genetic parameter estimation using mixed models, Bayesian approaches, and machine learning techniques. We evaluated these approaches using simulated and real breeding data sets. Our contribution in this study is a five-fold demonstration: (i) collecting data using an intermediate number of categories (1 to 3 and 1 to 5 scores) is the best strategy, even considering errors and subjectivity associated with visual scores; (ii) in the context of genomic prediction, Linear Mixed Models and Bayesian Linear Regression Models are robust to the normality violation, but marginal gains can be achieved when using Bayesian Ordinal Regression Models (BORM) and Random Forest Classification technique; (iii) genetic parameters are better estimated using BORM; (iv) our conclusions using simulated data are also applicable to real data in autotetraploid blueberry, which can guide breeders decisions; and (v) a comparison of continuous and categorical phenotype testing for complex traits with low heritability, found that investing in the evaluation of 600-1000 categorical data points with low error, when it is not feasible to collect continuous phenotypes, is a strategy for improving predictive abilities. Our findings suggest the best approaches for effectively using categorical traits to explore genetic information in breeding programs, and highlight the importance of investing in the training of evaluator teams and in high-quality phenotyping. Key messageAn approach for handling categorical data with potential errors and subjectivity in scores was evaluated in simulated and blueberry recurrent selection breeding schemes to assist breeders in their decision-making.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification and Regression Models for Genomic Selection of Skewed Phenotypes: A Case for Disease Resistance in Winter Wheat (Triticum aestivum L.) 97%
- Using local convolutional neural networks for genomic prediction 97%
- Harnessing genetic diversity in the USDA pea (Pisum sativum L.) germplasm collection through genomic prediction 97%
Similar papers in this journal
- Improved Genomic Prediction Performance with Ensembles of Diverse Models 97%
- PyBrOpS: a Python package for breeding program simulation and optimization for multi-objective breeding 96%
- Back to the future 2: the implications of germplasm structure on thebalance between short and long-term genetic gain in a changingtarget population of environments 96%
Similar papers in this journal
- Comparison of Genomic Selection Models for Exploring Predictive Ability of Complex Traits in Breeding Programs 97%
- Multi-Trait Machine and Deep Learning Models for Genomic Selection using Spectral Information in a Wheat Breeding Program 97%
- Optimal implementation of genomic selection in clone breeding programs - exemplified in potato: I. Effect of selection strategy, implementation stage, and selection intensity on short-term genetic gain 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.