Sex Checking by Zygosity Distributions
Molina-Sedano, O.; Mas Montserrat, D.; Ioannidis, A. G.
Show abstract
MotivationIn genomic and clinical studies, verifying concordance between self-reported and genotype-inferred sex is a crucial quality control step, since mismatches arising from mislabeling or aneuploidies can bias downstream analyses and affect diagnostic accuracy. Existing approaches typically require substantial auxiliary data, and often require manual threshold tuning. There remains a need for a streamlined, reference-free method that generalizes across different data modalities--including whole-genome, single-sample and array--without requiring additional files or parameter tuning. ResultsWe present Zigo, a novel ML-based sex-checking method that operates solely on a standard VCF file, designed using X-chromosome genotype class distributions across sexes. Our model was trained on synthetic data incorporating standard demographic models and empirical recombination maps to ensure realistic genetic architecture and population structure. We simulate WGS, array, and single-sample files for broad applicability. Unlike traditional methods, we eliminate manual thresholding by distilling learned discriminative patterns into a single polynomial equation that determines genetic sex directly from normalized genotype counts. We validated Zigo on independent datasets, including 1000 Genomes, UK Biobank, and HGDP. Additional experiments assessed robustness under reduced variant availability through random SNP subsampling and allele-frequency filtering. Across all evaluations, the model achieved state-of-the-art accuracy, high time efficiency, and strong generalization, even with severely limited variant sets. AvailabilityWe release our sex-checking tool as an open-source command-line interface (CLI) under GitHub at https://github.com/AI-sandbox/zigo. Supplementary informationSupplementary data are available in the Supplementary Material at the end of this article.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep convolutional and conditional neural networks for large-scale genomic data generation 95%
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 95%
- Ancestral Haplotype Reconstruction in Endogamous Populations using Identity-By-Descent 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.