Revealing the range of maximum likelihood estimates in the admixture model.
Heinzel, C. S.; Baumdicker, F.; Pfafffelhuber, P.
Show abstract
Many ancestry inference tools, including STRUCTURE and ADMIXTURE, rely on the admixture model to infer both, allele frequencies p and individual admixture proportions q for a collection of individuals relative to a set of hypothetical ancestral populations. We show that under realistic conditions the likelihood in the admixture model is typically flat in some direction around a maximum likelihood estimate [Formula]. In particular, the maximum likelihood estimator is non-unique and there is a complete spectrum of possible estimates. Common inference tools typically identify only a few points within this spectrum. We provide an algorithm which computes the set of equally likely [Formula], when starting from [Formula]. It is analytic for K = 2 ancestral populations and numeric for K > 2. We apply our algorithm to data from the 1000 genomes project, and show that inter-European estimators of q can come with a large set of equally likely possibilities. In general, markers with large allele frequency differences between populations in combination with individuals with concentrated admixture proportions lead to small areas with a flat likelihood. Our findings imply that care must be taken when interpreting results from STRUCTURE and ADMIXTURE if populations are not separated well enough.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimating the time since admixture from phased and unphased molecular data 97%
- FSTruct: An Fst-based tool for measuring ancestry variation in inference of population structure 96%
- Inferring the timing and strength of natural selection and gene migration in the evolution of chicken from ancient DNA data 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.