Why Accuracy Metrics Fall Short in Comparing Phenomic and Genomic Prediction Models
Wang, F.; Feldmann, M. J.; Runcie, D. E.
Show abstract
Phenomic selection is a new paradigm in plant breeding that uses high-throughput phenotyping technologies and machine learning models to predict traits of new individuals and make selections. This can allow breeders to evaluate more plants in higher throughput more accurately, resulting in faster rates of gain and reduced labor costs. However, phenomic prediction models are frequently benchmarked against genomic prediction models using cross-validation to demonstrate their usefulness to breeders. We argue that this is inappropriate for two reasons: 1) differences in the accuracy statistic measured by cross-validation do not reliably indicate differences in the accuracy parameter of the breeders equation, which we show analytically and through re-analysis of data from three representative phenomic prediction studies, and 2) phenomic and genomic selection tools influence other parameters of the breeders equation, so comparing accuracy, even if done properly, is insufficient to advocate for one approach over the other. We conclude that phenomic selection may be useful, but comparisons of accuracy between genomic prediction and phenomic prediction models are not.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Back to the future 2: the implications of germplasm structure on thebalance between short and long-term genetic gain in a changingtarget population of environments 97%
- Improved Genomic Prediction Performance with Ensembles of Diverse Models 96%
- PyBrOpS: a Python package for breeding program simulation and optimization for multi-objective breeding 96%
Similar papers in this journal
- Reciprocal recurrent selection based on genetic complementation: An efficient way to build heterosis in diploids due to directional dominance 96%
- Aerial High-Throughput Phenotyping Enabling Indirect Selection for Grain Yield at the Early-generation Seed-limited Stages in Breeding Programs 95%
- Phasing and imputation of single nucleotide polymorphism data of missing parents of bi-parental plant populations 95%
Similar papers in this journal
- Multi-trait multi-environment genomic prediction of preliminary yield trials in pulse crops 96%
- Genomic prediction of stalk lodging resistance and the associated intermediate phenotypes in maize using whole-genome resequence and multi-environmental data 96%
- Comparison of Genomic Selection Models for Exploring Predictive Ability of Complex Traits in Breeding Programs 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.