Efficient Estimation of Nucleotide Diversity and Divergence using Depth Information
Mirchandani, C. D.; Enbody, E.; Sackton, T. B.; Corbett-Detig, R.
Show abstract
The increasing scale of population genomic datasets presents computational challenges in estimating summary statistics such as nucleotide diversity ({pi}) and divergence (dxy). Unbiased estimates of diversity require knowledge of missing data and existing tools require all-sites VCFs. However, generating these files is computationally expensive for large datasets. Here, we introduce Callable Loci And More (clam), a tool that leverages callable loci--determined from depth information--to estimate population genetic statistics using a variant-only VCF. This approach offers improvements in storage footprint and computational performance compared to contemporary methods. We benchmark clam using a large muskox dataset and demonstrate that it produces unbiased estimates of {pi} while reducing runtime and storage requirements, compared to an existing approach. clam provides an efficient and scalable alternative for population genomic analyses, facilitating the study of increasingly large and diverse datasets. clam is available as a standalone program and integrated into snpArcher for efficient reproducible population genomic analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Estimating allele frequencies, ancestry proportions and genotype likelihoods in the presence of mapping bias 94%
- WeavePop: A bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations 94%
- kGWASflow: a modular, flexible, and reproducible Snakemake workflow for k-mers-based GWAS 94%
Similar papers in this journal
- The Practical Haplotype Graph, a platform for storing and using pangenomes for imputation 95%
- AlphaFamImpute: high accuracy imputation in full-sib families from genotype-by-sequencing data 95%
- Polaris: Polarization of ancestral and derived polymorphic alleles for inferences of extended haplotype homozygosity in human populations. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.