Back

Modelling complex traits with ancestral recombination graphs

Lee, H.; Pope, N. S.; Kelleher, J.; Gorjanc, G.; Ralph, P. L.

2025-07-18 genomics
10.1101/2025.07.14.664631 bioRxiv
Show abstract

Ancestral recombination graphs (ARGs) are an attractive means for quantitative genetic analysis of complex traits because they encode the realized genetic relatedness between a sample of individuals in the presence of genetic drift, recombination, and mutation. Data structures for efficiently storing ARGs can also be used to rapidly process millions of genomes, and are thus promising for fitting linear mixed models (LMMs) to large phenotype and genome datasets. Here, we study the problems of variance component estimation and prediction of genetic values with ARGs, by describing a generative model of complex traits with additive effects on an ARG, and then developing algorithms that use the ARG to solve these problems efficiently on biobank-scale datasets. We observe nearly linear scaling of runtime with sample size, which is achieved by using the succinct tree sequence representation of the ARG for implicit matrix-vector products, along with modern randomized linear algebra algorithms. We estimate variance components using restricted maximum likelihood (REML), which we find performs substantially better than the Haseman- Elston method. In simulation tests, both variance component estimation and prediction of genetic values (using the best linear unbiased predictor, BLUP) perform nearly as well with inferred ARGs as with true ARGs. We also discuss interpretations of the variance component estimates as mutational variance and additive genetic variance. We provide an implementation of the algorithms as a Python package tslmm, which leverages the tree sequence library tskit.

Published in GENETICS (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.