Back

Coalescent-Based Time-Stratified Statistics Reveal Population Structure Dynamics using the Ancestral Recombination Graph

Deng, Y.; Pritchard, J. K.; Spence, J. P.

2026-08-18 evolutionary biology
10.64898/2026.08.11.744210 bioRxiv
Show abstract

Many questions in population genetics are concerned with reconstructing evolutionary history through time, such as inferring how population structure has changed throughout the past. Yet, many existing approaches have only an implicit temporal component, using quantities such as allele frequency or haplotype length as rough proxies for age. Recent advances in the inference of Ancestral Recombination Graphs (ARGs) have made it possible to estimate the entire sequence of local genealogies along the genome. These genealogies explicitly encode how samples are related to each other at different time points in the past, enabling the inference of how population structure has changed over time. To this end, recent work has used ARGs to define time-stratified versions of widely-used population genetics summary statistics in an attempt to capture the population structure present within a particular time window. Here, we show that naive approaches result in statistics that cannot be interpreted solely in terms of the population structure present within the time window they are targeting. To address this problem, we introduce a framework of coalescent-based time-stratified statistics, which use coalescence probabilities to partition classical summary statistics into interval-specific contributions. Using coalescent simulations, we demonstrate that these statistics accurately isolate population structure at different temporal depths and avoid spurious signals. Our results highlight the necessity of integrating coalescent theory into ARG-based temporal analyses and provide a principled and practical foundation for studying the dynamics of population structure through time.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.