A novel hyper-parameter can increase the prediction accuracy in a single-step genetic evaluation
Neshat, M.; Lee, S.; Momin, M. M.; Truong, B.; van der Werf, J. H. J.; Lee, S. H.
Show abstract
The H-matrix best linear unbiased prediction (HBLUP) method has been widely used in livestock breeding programs. It can integrate all information, including pedigree, genotypes, and phenotypes on both genotyped and non-genotyped individuals into one single evaluation that can provide reliable predictions of breeding values. The existing HBLUP method (e.g., that implemented in BLUPf90 software) requires hyper-parameters that should be adequately optimised as otherwise the genomic prediction accuracy may decrease. In this study, we assess the performance of HBLUP using various hyper-parameters such as blending, tuning and scale factor in simulated as well as real data on Hanwoo cattle. In both simulated and cattle data, we show that blending is not necessary, indicating that the prediction accuracy decreases when using a blending hyper-parameter < 1. The tuning process (adjusting genomic relationships accounting for base allele frequencies) improves prediction accuracy in the simulated data, confirming previous studies, although the improvement is not statistically significant in the Hanwoo cattle data. We also demonstrate that a scale factor, , which determines the relationship between allele frequency and per-allele effect size, can improve the HBLUP accuracy in both simulated and real data. Our findings suggest that an optimal scale factor should be considered to increase the prediction accuracy, in addition to blending and tuning processes, when using HBLUP. Author SummaryDespite significant advancements in genotyping technologies, the capability to predict the phenotypes of complex traits is still limited. H-matrix best linear unbiased prediction (HBLUP) method has been used to tackle this limitation to demonstrate a promising prediction accuracy. However, the performance of HBLUP depends heavily on the optimisation of hyper-parameters (e.g. blending and tuning). In this study, we introduce a scale factor (), as a new hyper-parameter in HBLUP, which accounts for the relationship between allele frequency and per-allele effect size. Using simulation and real data analysis, we investigate the impact of the hyper-parameters (blending, tuning, and scale factor) on the performance of HBLUP. In general, the blending process may not improve the prediction accuracy for simulation and cattle data although a marginally improved prediction accuracy is observed with a blending hyper-parameter = 0.86 for one of carcass traits in the cattle data. In contrast, the tuning process can increase the HBLUP accuracy particularly in simulated data. Furthermore, we observe that an optimal scale factor plays a significant role in improving the prediction accuracy in both simulated and real data, and the improvement is relatively large compared with blending and tuning processes. In this context, we propose considering the scale factor as a hyper-parameter to increase the predictive performance of HBLUP.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dimensionality of genomic information and its impact on GWA and variant selection: a simulation study 97%
- Optimisation of the core subset for the APY approximation of genomic relationships 96%
- Bayesian genomic models boost prediction accuracy for resistance against Streptococcus agalactiae in Nile tilapia (Oreochromus nilioticus) 95%
Similar papers in this journal
- Genomic prediction using machine learning: A comparison of the performance of regularized regression, ensemble, instance-based and deep learning methods on synthetic and empirical data 94%
- Preselection of QTL markers enhances accuracy of genomic selection in Norway spruce 93%
- Robust estimation of heritability and predictive accuracy in plant breeding: evaluation using simulation and empirical data 93%
Similar papers in this journal
Similar papers in this journal
- Multifactorial Methods Integrating Haplotype and Epistasis Effects for Genomic Estimation and Prediction of Quantitative Traits 97%
- Validation of the linear regression method to evaluate population accuracy and bias of predictions for non-linear models 94%
- Optimizing Sequencing Resources in Genotyped Livestock Populations Using Linear Programming 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.