Estimating time-varying selection coefficients from time series data of allele frequencies
Mathieson, I.
Show abstract
Time series data of allele frequencies are a powerful resource for detecting and classifying natural and artificial selection. Ancient DNA now allows us to observe these trajectories in natural populations of long-lived species such as humans. Here, we develop a hidden Markov model to infer selection coefficients that vary over time. We show through simulations that our approach can accurately estimate both selection coefficients and the timing of changes in selection. Finally, we analyze some of the strongest signals of selection in the human genome using ancient DNA. We show that the European lactase persistence mutation was selected over the past 5,000 years with a selection coefficient of 2-2.5% in Britain, Central Europe and Iberia, but not Italy. In northern East Asia, selection at the ADH1B locus associated with alcohol metabolism intensified around 4,000 years ago, approximately coinciding with the introduction of rice-based agriculture. Finally, a derived allele at the FADS locus was selected in parallel in both Europe and East Asia, as previously hypothesized. Our approach is broadly applicable to both natural and experimental evolution data and shows how time series data can be used to resolve fine-scale details of selection.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Flexible mixture model approaches that accommodate footprint size variability for robust detection of balancing selection 97%
- Fast and accurate estimation of selection coefficients and allele histories from ancient and modern DNA 97%
- Disentangling signatures of selection before and after European colonization in Latin Americans 97%
Similar papers in this journal
Similar papers in this journal
- Learning the properties of adaptive regions with functional data analysis 98%
- A novel expectation-maximization approach to infer general diploid selection from time-series genetic data 98%
- An approximate full-likelihood method for inferring selection and allele frequency trajectories from DNA sequence data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.