Calculating and interpreting FST in the genomics era
de Jong, M. J.; van Oosterhout, C.; Hoelzel, R.; Janke, A.
Show abstract
The relative genetic distance between populations is commonly measured using the fixation index (FST). Traditionally inferred from allele frequency differences, the question arises how FST can be estimated and interpreted when analysing genomic datasets with low sample sizes. Here, we advocate an elegant solution first put forward by Hudson et al. (1992): FST = (Dxy -{pi} xy)/Dxy, where Dxy and{pi} xy denote mean sequence dissimilarity between and within populations, respectively. This multi-locus FST-metric can be derived from allele frequency data, but also from sequence alignment data alone, even when sample sizes are low and/or unequal. As with other FST-metrices, the numerator denotes net divergence (Da), which is equivalent to the f2-statistic and Neis D (for realistic estimates of Dxy and{pi} xy). In terms of demographic inference, net divergence measures the difference in increase of Dxy and{pi} xy since the population split, owing to a reduction of coalescence times within populations as a result of genetic drift. Because different combinations of{Delta} Dxy and{Delta}{pi} xy can produce identical FST-estimates, no universal relationship exists between FST and population split time. Still, in case of recent population splits, when novel mutations are negligible, FST-estimates can be accurately converted into coalescent units ({tau}. i.e., split time in multiples of 2Ne). This then allows to quantify gene tree discordance, without the need for multispecies coalescent based analyses, using the formula: Pdiscordance = [2/3]{middle dot}(1 - FST). To facilitate the use of the Hudson FST-metric, we implemented new utilities in the R package SambaR.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.