Modeling recent positive selection in Americans of European ancestry
Temple, S. D.; Waples, R. D.; Browning, S. R.
Show abstract
Recent positive selection can result in an excess of long identity-by-descent (IBD) haplotype segments. The statistical methods that we propose here address three major objectives in studying selective sweeps: scanning for regions of interest, identifying possible sweeping alleles, and estimating a selection coefficient s. First, we implement a selection scan to locate regions of excess IBD rate. Second, we develop a statistic to rank alleles that are in strong linkage disequilibrium with a putative sweeping allele. We aggregate these scores to estimate the allele frequency of the sweeping allele, even if it is not genotyped. Third, we propose an estimator for the selection coefficient and quantify uncertainty using the parametric bootstrap. Comparing against state-of-the-art methods in extensive simulations, we show that our methods are better at identifying sweeping alleles that are at low frequency and at estimating S when S [≥] 0.015. We apply these methods to study positive selection in European ancestry samples from the TOPMed project. We analyze eight loci where the IBD rate is more than four standard deviations above the population median. The IBD rate at LCT is thirty-five standard deviations above the population median, and our estimates of its selection coefficient imply strong selection within the past two hundred generations. Overall, we present robust and accurate approaches to study very recent adaptive evolution without knowing the identity of the causal allele or using time series data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Reconstructing the history of founder events using genome-wide patterns of allele sharing across individuals 98%
- Regularized sequence-context mutational trees capture variation in mutation rates across the human genome 96%
- Learning the properties of adaptive regions with functional data analysis 96%
Similar papers in this journal
- Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies 97%
- A resource-efficient tool for mixed model association analysis of large-scale data 96%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.