Technical artifact drives apparent deviation from Hardy-Weinberg equilibrium at CCR5-Δ32 and other variants in gnomAD
Karczewski, K. J.; Gauthier, L. D.; Daly, M. J.
Show abstract
Following an earlier report suggesting increased mortality due to homozygosity at the CCR5-{Delta}32 allele1, Wei and Nielsen recently suggested a deviation from Hardy-Weinberg Equilibrium (HWE) observed in public variant databases as additional supporting evidence for this hypothesis2. Here, we present a re-analysis of the primary data underlying this variant database and identify a pervasive genotyping artifact, especially present at long insertion and deletion polymorphisms. Specifically, very low levels of contamination can affect the variant calling likelihood models, leading to the misidentification of homozygous individuals as heterozygous, and thereby creating an apparent depletion of homozygous calls, which is especially prominent at large insertions and deletions. The deviation from HWE observed at CCR5-{Delta}32 is a consequence of this specific genotyping error mode rather than a signature of selective pressure at this locus.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discordant calls across genotype discovery approaches elucidate variants with systematic errors 96%
- Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays 94%
- Mitochondrial DNA variation across 56,434 individuals in gnomAD 94%
Similar papers in this journal
- Inferring compound heterozygosity from large-scale exome sequencing data 94%
- Long read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits 94%
- Adjusting for Common Variant Polygenic Scores Improves Yield in Rare Variant Association Analyses 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.