Modelling complex population structure using F-statistics and Principal Component Analysis
Peter, B. M.
Show abstract
Human genetic diversity is shaped by our complex history. Data-driven methods such as Principal Component Analysis (PCA) are an important population genetic tool to understand this method. Here, I contrast PCA with a set of statistics motivated by trees (F-statistics). Here, I show that these two methods are closely related, and I derive explicit connections between the two approaches. I show that F-statistics have a simple geometrical interpretation in the context of PCA, and that orthogonal projections are the key concept to establish this link. I illustrate my results on two examples, one of local, and one of global human diversity. In both examples, I find that just using the first few PCs provides good population structure is sparse, and only a few components contribute to most statistics. Based on these results, I develop novel visualizations that allow for investigating specific hypotheses, checking the assumptions of more sophisticated models. My results extend F-statistics to non-discrete populations, moving towards more complete and less biased descriptions of human genetic variation.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- SelNeTime: a python package inferring effective population size and selection intensity from genomic time series data 93%
- Exploring the effects of ecological parameters on the spatial structure of genealogies 93%
- An approximate likelihood method reveals ancient gene flow between human, chimpanzee and gorilla 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.