Back

Expanded unbiased population-genomic summary statistics in pixy

Samuk, K.; Stone, M.; McAuley, E.; Bailey, N. P.; Lucas, G.; Plaza, M.

2026-07-30 bioinformatics
10.64898/2026.07.27.741083 bioRxiv
Show abstract

Estimates of population-genetic summary statistics are often computed from variant-only VCFs, which commonly omit invariant sites (i.e. sites with only homozygous reference genotypes across all samples). However, these sites are critical for correctly estimating per site statistics such as nucleotide diversity ({pi}) and between population divergence (dxy). Our software pixy addressed this issue by providing support for computing statistics directly from "all-sites" VCFs that encode invariant positions explicitly (Korunes and Samuk 2021). pixy has since been widely adopted and used in a large variety of population genetic studies. Here we present a major update to pixy, which expands the original tool in four major areas. First, we provide implementations of new estimators of Wattersons {theta} and Tajimas D that are unbiased with respect to missing data. Secondly, we provide support for arbitrary ploidy and multiallelic sites. Third, we introduce a variety of optimizations, including multicore execution and a roughly order-of-magnitude reduction in per-worker memory footprint. Finally, we have broadly modernized our code base, test suite, and community contribution pathway. We validate these new features on simulated polyploid and multiallelic datasets, as well as two empirical datasets (one diploid and one autotetraploid), and benchmark the new multicore scaling. We position pixy relative to contemporary tools and discuss shared limitations of VCF-based estimators. By integrating support for diverse ploidies, allelic complexity, and genome-wide scale in one workflow, pixy provides the population genetics community with a straightforward tool for estimates of key summary statistics that are unbiased with respect to missing data. pixy remains completely open-source (MIT-licensed) and easily installable with the conda package management system.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.