Optimizing genomic selection: A comparison of SNP selection strategies for reduced-density panels in beef cattle
Ogunbawo, A. R.; Mulim, H. A.; Hidalgo, J.; Ventura, H. T.; Souza, N. O.; Oliveira, H. R.
Show abstract
The exponential increase in the number of genotyped animals, combined with the availability of high-density SNP chips has introduced computational challenges for routine genomic evaluations, particularly during the construction of the genomic relationship matrix. Although higher-density SNP panels can facilitate the identification of causal mutations, their use substantially increases computational requirements without a proportional gain in genomic prediction performance. To optimize computational efficiency while maintaining accuracy of genomic predictions, this study compared five SNP selection strategies (i.e., random sampling, random sampling with inclusion of informative SNPs, linkage disequilibrium (LD)-based pruning, a Shannon entropy-based machine learning approach, and [[EQUATION]]-based prioritization) to develop reduced-density panels for Nellore cattle. Using high-density (HD) genotype data comprising 437,650 SNPs from 304,782 animals (after quality control) as reference, three reduced-density panels (25K, 45K, and 65K SNPs) panels were tested across five traits (i.e., Age at first calving, Stayability, Weaning weight, Yearling weight, Muscling) with diverse genetic architectures. Genomic estimated breeding values (GEBVs) derived from these reduced panels were compared to those obtained from the HD reference panel using Pearsons correlations, under both genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP) methods. In the GBLUP model, prediction accuracy generally improved with increased marker density. Random selection with and without the informative SNPs consistently yielded the highest accuracies, whereas the [[EQUATION]]-based approach showed the lowest agreement with the HD reference across all densities. In contrast, ssGBLUP demonstrated strong robustness to marker reduction, producing uniformly high correlations {approx}1.00) across all SNP densities and selection strategies. These findings indicate that optimized low-density SNP panels maintain prediction accuracy comparable to HD panels, offering a cost-effective tool for large-scale genomic evaluations.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Across-breed analyses of genome-wide association studies for stature and mammary gland morphology in cattle reveal pleiotropic effects of the Friesian POLLED haplotype 95%
- Back to the horns: a reconstruction of the ancestral horn state through distinct types of recombination events. 94%
- Spatial modelling improves genomic evaluation in Tanzanian smallholder admixed dairy cattle 94%
Similar papers in this journal
- Signature of selection in composite Vrindavani cattle of India 94%
- Variance component estimates, phenotypic characterization, and genetic evaluation of bovine congestive heart failure in commercial feeder cattle 94%
- Deciphering cattle temperament measures derived from a four-platform standing scale using genetic factor analytic modeling 93%
Similar papers in this journal
- Multi-trait meta-analyses reveal 25 quantitative trait loci for economically important traits in Brown Swiss cattle 95%
- GWAS and Fine-Mapping of Livability and Six Disease Traits in Holstein Cattle 94%
- Allelic Variation in CYP3A4 and PLB1 Drives Feed Efficiency and Immunometabolic Resilience in Beef Cattle 94%
Similar papers in this journal
- Extensive genome-wide association analyses identify genotype-by-environment interactions of growth traits in Simmental cattle 92%
- BoLA-DRB3 gene haplotypes show divergence in native Sudanese cattle from Taurine and Zebu breeds 92%
- A reduced SNP panel optimised for non-invasive genetic assessment of a genetically impoverished conservation icon, the European bison 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.