Best practices for improving alignment and variant calling on human sex chromosomes
Oill, A. M. T.; Plaisier, S. B.; Phung, T. N.; Wilson, M. A.
Show abstract
Sex chromosome complement is the largest karyotypic variation observed in humans. X and Y chromosomes were once a pair of homologous autosomes. Although chromosome X and Y differentiated from one another, they still share high levels of sequence similarity in some regions, like the pseudoautosomal regions (PARs) and the X-transposed region (XTR). The sex chromosomes violate some assumptions of autosomal pairs, but are not always processed separately in genomics analyses. Here, we undertook a simulation study to assess the effects of standard autosomal versus sex chromosome complement-informed alignment, variant calling, and variant filtering strategies on variants detected on the human sex chromosomes. We find that aligning samples to a reference genome informed by the sex chromosome complement of the sample increases the number of true positives called in the PARs, and, in XX-samples only, also the XTR. In contrast, in XY-samples, masking the XTR during alignment results in a ten-fold higher rate of false positives. We further find that haploid calling on the sex chromosomes in XY-samples reduces the number of false positives compared to diploid calling, but does not decrease the number of false negatives. Improving the accuracy of variant calling results in detection of variants that could be relevant to studies of health and disease, including variants we recovered in genes implicated in cardiomyopathy, immunodeficiency, and Alzheirmers disease, among others. We recommend future genomic analyses implement the following best practices for detecting variants: aligning samples to versions of the human reference genome informed by the sex chromosome complement of the sample and using accurate ploidy parameters when calling variants.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Accurate, ultra-low coverage genome reconstruction and association studies in Hybrid Swarm mapping populations 94%
- A Simple Deep Learning Approach for Detecting Duplications and Deletions in Next-Generation Sequencing Data 94%
- Locating the sex determining region of linkage group 12 of guppy (Poecilia reticulata) 94%
Similar papers in this journal
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 98%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 94%
- Protein Coding Variation In Outbred Laboratory Mouse Stocks Provides A Molecular Basis For Distinct Research Applications 93%
Similar papers in this journal
- Variation in mutation, recombination, and transposition rates in Drosophila melanogaster and Drosophila simulans 94%
- SVCollector: Optimized sample selection for cost-efficient long-read population sequencing 94%
- Assessing and mitigating privacy risk of sparse, noisy genotypes by local alignment to haplotype databases 94%
Similar papers in this journal
- Major sex differences in allele frequencies for X chromosome variants in the 1000 Genomes Project data 95%
- Estimating indirect parental genetic effects on offspring phenotypes using virtual parental genotypes derived from sibling and half sibling pairs 94%
- A robust and adaptive framework for interaction testing in quantitative traits between multiple genetic loci and exposure variables 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.