Back

Predicting phenotypes with one step genetic decision trees

Blommaert, J.; Bayer, P. E.; Ashton, D. T.; Samuels, G.; Jesson, L.; Wellenreuther, M.

2026-07-19 genomics
10.1101/2025.05.29.656727 bioRxiv
Show abstract

Genomic prediction of complex traits is limited when phenotype records are restricted and when using linear models. Increasing the amount of phenotypic data with high-throughput, image-based phenotyping could result in better genomic prediction and stronger signals in variant detection. Here, we analysed phenotypic and genomic data from a selectively bred cohort of the Australasian snapper (Chrysophrys auratus) to identify genetic variants associated with growth traits. We used a high-throughput phenotyping pipeline to extract 13 measurements of size from images. Phenotypic correlations among image-derived and manually measured traits (weight, fork length), together with heritabilities, were analysed. All measurements were significantly positively correlated with each other, and heritability ranged from 0.20-0.38. Genome-wide association studies (GWAS) identified 28 growth-associated SNPs, while GBLUP was used to predict phenotypes, and XGBoost machine-learning models were used to jointly predict phenotypes and report important variants. Both GBLUP (mean R2 = 0.50) and XGBoost (mean R2 = 0.77) performed well on the training data, but performance dropped on testing sets (both = 0.11), which decreased further when accounting for genetic relatedness (both = 0.06). Despite this, approximately 20% of the genetic variance for growth traits was captured by the models, and feature importance from XGBoost reflected signals seen in GWAS. Our findings highlight the utility of integrating computer vision-based phenotyping with GWAS, GBLUP, and ML for trait prediction. Despite detecting shared biological signals as GWAS, genomic prediction faces challenges with population structure and relatedness that are inherent in breeding programmes of mass spawning species, including many aquatic species. Article summaryIncorporating genetic information into selective breeding programmes can accelerate gains but may also miss gene interactions in complex traits. Machine learning approaches, such as decision trees, can capture those relationships and potentially improve genomic predictions. We used high-throughput computer vision phenotyping to uncover biological signal for genetic growth variants involved in the Australasian snapper. Both types of models captured 18-40% of the genetic variation, but the prediction accuracies were hampered by the population structure in this cohort of mass-spawning fish.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.