Assessing Hardy-Weinberg Equilibrium in T2T-aligned 1000 Genomes Project
Garg, E.; Romain, J.; Sun, L.; Paterson, A. D.
Show abstract
Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. We examined genome-wide sex-specific HWE deviations across populations in the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequence of the 1000 Genomes Project data of 2,490 individuals. Our analysis was restricted to bi-allelic SNPs with non-missing genotypes and MAF>=5% in both sexes of the five super-populations. We employed a robust allele-based approach for HWE testing, which enabled the quantification of directional deviations from HWE. A second-order omnibus meta-analysis combining results from the five super-populations and both sexes revealed that 0.9% autosomal SNPs exhibited a significant deviation from HWE at p<5e-8. Most of these deviations were found to be associated with genomic features relating to poor sequence quality. Filtering results to reliable genomic regions yielded 255 autosomal and 1 NPR X chromosomal SNPs, of which 140 autosomal SNPs also showed significant heterogeneity across populations but not across sexes. 8 SNPs in a 15-bp region on chr14 showed excess heterozygosity in both sexes of the AFR (African) super-population. We also generated a well-performing multivariate predictor of HWD (deviation from HWE) using multiple sequence features, which could be combined with HWD estimates in future studies to select SNPs that deviate from HWE due to technical rather than biological reasons. Author SummaryWe conducted a specific quality control test, which compares the observed and expected genotype counts, on an updated version of the 1000 Genomes Project whole genome sequence data generated on [~]2500 individuals. We first performed this analysis by grouping the data by ancestry and sex. We then combined and contrasted the group results. We found that most regions that differed between observed and expected counts overlapped regions of the genome which are difficult to sequence using current short read technology. In the remaining regions we found an interesting cluster of SNPs in a single ancestry, where there is a gross excess of heterozygous genotypes. GWASes typically use a standard strict threshold for this quality control test for genotyping arrays to remove SNPs. Here we suggest a more nuanced approach that is applicable to whole genome sequence data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Major sex differences in allele frequencies for X chromosome variants in the 1000 Genomes Project data 98%
- Increased ultra-rare variant load in an isolated Scottish population impacts exonic and regulatory regions 95%
- Adjusting for principal components can induce spurious associations in genome-wide association studies in admixed populations 94%
Similar papers in this journal
- Evaluation of imputation performance of multiple reference panels in a Pakistani population 96%
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 95%
- A reference panel for linkage disequilibrium and genotype imputation using whole-genome sequencing data from 2,680 participants across India 94%
Similar papers in this journal
- Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays 95%
- Aberrant landscapes of maternal meiotic crossovers contribute to aneuploidies in human embryos 93%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.