Back

Inferring drift, genetic differentiation, and admixture graphs from low-depth sequencing data

Rasmussen, M. S.; Wiuf, C.; Albrechtsen, A.

2024-01-31 bioinformatics
10.1101/2024.01.29.577762 bioRxiv
Show abstract

A number of popular methods for inferring the evolutionary relationship between populations require essentially two components: First, they require estimates of f2-statistics, or some quantity that is a linear combination of these. Second, they require estimates of the variability of the statistic in question. Examples of methods in this class include qpGraph and TreeMix. It is known, however, that these statistics are biased when based on genotype calls at low depth. Moreover, as we show, this leads to downstream inference of significantly distorted trees. To solve this problem, we demonstrate how to accurately and efficiently compute a broad class of statistics from low-depth whole-genome sequencing data, including estimates of their standard errors, by using the site frequency spectrum. In particular, we focus on f2 and the sample covariance of allele frequencies to show how this method leads to accurate estimate of drift when fitting trees using qpGraph and TreeMix with low-depth data. However, the same considerations lead to uncertainty estimates for a variety of other statistics, including heterozygosity, kinship estimates (e.g. King), and quantities relating to genetic differentiation such as Fst and Dxy.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.