Back

A New Likelihood-based Test for Natural Selection

Simon, H. E.; Huttley, G. A.

2021-07-04 genetics
10.1101/2021.07.04.451068 bioRxiv
Show abstract

We present a new statistic for testing for neutral evolution from allele frequency data summarised as a site frequency spectrum, which we call the relative likelihood neutrality test or{rho} . Classical methods of testing for natural selection, such as Tajimas D and its relatives, require the null model to have constant population size over time and therefore can confound demographic change with natural selection.{rho} can directly incorporate a null hypothesis reflecting general demographic histories. It has a natural Bayesian interpretation as an approximation to the log-probability of the null model, given the data. We use simulations to show that{rho} has greater power than Tajimas D to detect departure from neutrality for a range of scenarios of positive and negative selection. We also show how{rho} can be adapted to account for sequencing error. Application to the ACKR1 (FYO) gene in humans supported previous studies inferring positive selection in sub-Saharan populations which were based on inter-population comparisons. However, we did not find the signal of selection to be maximal in the region of the FY*O or Duffy-null allele in these populations. We also applied{rho} to investigate in greater detail a region on the 2q11.1 band of the human genome that has previously been identified as showing evidence of selection. This was done for a range of populations: for the European populations we incorporated a demographic history with a bottleneck corresponding to the putative out of Africa event. We were able to localise signals of selection to some specific regions and genes. Overall, we suggest that{rho} will be a useful tool for identifying genomic regions that may be subject to natural selection.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.