Back

SelNeTime: a python package inferring effective population size and selection intensity from genomic time series data

Uhl, M.; Bunel, P.; de Navascues, M.; boitard, s.; Servin, B.

2025-07-01 evolutionary biology
10.1101/2024.11.06.622284 bioRxiv
Show abstract

Genomic samples collected from a single population over several generations provide direct access to the genetic diversity changes occurring within a specific time period. This provides information about both demographic and adaptive processes acting on the population during that period. A common approach to analyze such data is to model observed allele counts in finite samples using a Hidden Markov Model (HMM) where hidden states are true allele frequencies over time (i.e. a trajectory). The HMM framework allows one to compute the full likelihood of the data, while accounting both for the stochastic evolution of population allele frequencies along time and for the noise arising from sampling a limited number of individuals at possibly spread out generations. Several such HMM methods have been proposed so far, differing mainly in the way they model the transition probabilities of the Markov chain. Following Paris et al. (2019a), we consider here the Beta with Spikes approximation, which avoids the computational issues associated to the Wright-Fisher model while still including fixation probabilities, in contrast to other standard approximations of this model like the Gaussian or Beta distributions. To facilitate the analysis and exploitation of genomic time series data, we present an improved version of Paris et al. (2019a) s approach, denoted SelNeTime, whose computation time is drastically reduced and which accurately estimates effective population size (assuming no selection) or the selection intensity at each locus (given a previously estimated value of N). We also evaluate the performance of this method in realistic situations where selection is present and both demography and selection need to be inferred. SelNeTime is implemented in a user friendly python package, which can also easily simulate genomic time series data under a user-defined evolutionary model and sampling design.

Published in Peer Community Journal (predicted rank #2) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.