D-BLUP: a differentiable genomic BLUP model with learnable variance and marker weights
Guhlin, J. G.; Dearden, P.
Show abstract
1.Genomic best linear unbiased prediction (GBLUP) is widely used for genomic selection in livestock and crop breeding. There is growing interest in connecting machine learning with genomic breeding value prediction. Although the BLUP formula is itself mathematically differentiable, existing implementations do not expose a differentiable computational graph and cannot be trained end-to-end. Here, we present [D]-BLUP, a differentiable implementation of the genomic mixed model in JAX that makes the BLUP solve part of the training process rather than an external step. The variance ratio {lambda} and optional block-level kernel weights are treated as trainable parameters, allowing these components to be learned directly from prediction error while preserving the BLUP structure used in animal and plant breeding. This means BLUP can now fit naturally within gradient-based models common in machine learning without losing its interpretability. The method solves the familiar system (Gw + {lambda}I) u = y* with either an unweighted VanRaden genomic relationship matrix G or a block-weighted variant Gw, via automatic differentiation, allowing simultaneous training. On a public dataset, [D]-BLUP with a fixed VanRaden kernel reproduces rrBLUP estimated breeding values and predictive performance, while learning {lambda}, and optionally SNP block weights, from a mean squared error objective. [D]-BLUP preserves the structure of BLUP used in routine breeding programs but makes it differentiable, maintaining its interpretability while making it embeddable in machine learning pipelines.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enabling interpretable machine learning for biological data with reliability scores 94%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 93%
- Learning Genetic Perturbation Effects with Variational Causal Inference 93%
Similar papers in this journal
- PyMSQ: a python package for fast Mendelian sampling (co)variance and haplotype-based similarity in genomic selection 96%
- disperseNN2: a neural network for estimating dispersal distance from georeferenced polymorphism data 94%
- Haplotype based testing for a better understanding of the selective architecture 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.