A comprehensive representation of genotypic selection beyond two alleles, and its estimation from time-series data
Vellnow, N.; Gossmann, T. I.; Waxman, D.
Show abstract
Genetic diversity is central to the process of evolution. Both natural selection and random genetic drift are influenced by the level of genetic diversity of a population, since selection acts on diversity while drift samples from it. At a given locus in a diploid population, each individual carries only two alleles, but the population as a whole can possess a much larger number of alleles, with the upper limit constrained by twice the population size. This allows for many possible types of homozygotes and heterozygotes. Moreover, there are clinically significant loci, such as those related to the MHC complex, the ABO blood types, and cystic fibrosis, that exhibit a large number of alleles. Despite this, much of population genetic theory, and data analysis, are limited to biallelic loci, and up to now there is no flexible expression for the force of selection which applies for arbitrary numbers of alleles (and thus any number of heterozygotes), and which accommodates diverse fitness regimes. Here, we derive an expression, for the force of selection, which explicitly separates the effects of genetic diversity from the effects of fitness of different genotypes. The result presented facilitates our understanding and analysis of selection, and applies in a variety of different situations involving multiple alleles. This includes situations where fitnesses are additive, multiplicative, randomly fluctuating, frequency-dependent, and can involve explicit gene interactions, such as heterozygote advantage. We show how our results can be used to estimate fitness effects from allele frequency trajectories and employ such an approach on existing data in the literature on experimental yeast evolution. We uncover evidence of wide-spread heterozygote advantage across the yeast genome.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.