Deviations from Hardy Weinberg Equilibrium at CCR5-Δ32 in Large Sequencing Data Sets
Wei, X.; Nielsen, R.
Show abstract
Previous analyses of the UK Biobank (UKB) genotyping array data in the CCR5-{Delta}32 locus show evidence for deviations from Hardy-Weinberg Equilibrium (HWE) and an increased mortality rate of homozygous individuals, consistent with a recessive deleterious effect of the deletion mutation. We here examine if similar deviations from HWE can be observed in the newly released UKB Whole Exome Sequencing (WES) data and in the sequencing data of the Genome Aggregation Database (gnomAD). We also examine the reliability of the genotype calls in the UKB array data. The UKB genotyping array probe targeting CCR5-{Delta}32 (rs62625034) and the WES of {Delta}32 are strongly correlated (r2 = 0.97). This contrasts to tag SNPs of CCR5-{Delta}32 in the UKB which have high missing data rates and imputation errors rates. We also show that, while different data sets are subject to different biases, both the UKB-WES and the gnomAD data have a deficiency of homozygous CCR5-{Delta}32 individuals compared to the HWE expectation (combined P-value < 0.01), consistent with an increased mortality rate in homozygotes. Finally, we perform a survival analysis on data from parents of UKB volunteers, that, while underpowered, is also consistent with the original report of a deleterious effect of CCR5-{Delta}32 in the homozygous state.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FiMAP: A Fast Identity-by-Descent Mapping Test for Biobank-scale Cohorts 93%
- Modeling the length distribution of gene conversion tracts in humans from the UK Biobank sequence data 93%
- Beyond SNP Heritability: Polygenicity and Discoverability of Phenotypes Estimated with a Univariate Gaussian Mixture Model 93%
Similar papers in this journal
Similar papers in this journal
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 93%
- Probabilistic estimation of identity by descent segment endpoints and detection of recent selection 93%
- Estimating gene conversion rates from population data using multi-individual identity by descent 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.