Back

Near perfect identification of half sibling versus niece/nephew avuncular pairs without pedigree information or genotyped relatives

Sapin, E.; Keller, M. C.; Kelly, K.

2026-01-07 bioinformatics
10.64898/2026.01.06.697070 bioRxiv
Show abstract

MotivationLarge-scale genomic biobanks contain thousands of second-degree relatives with missing pedigree metadata. Accurately distinguishing half-sibling (H-S) from niece/nephew-avuncular (N/N-A) pairs--both sharing approximately 25% of the genome--remains a significant challenge. Current SNP-based methods rely on aggregate Identical-By-Descent (IBD) segment counts and age differences, but substantial distributional overlap leads to high misclassification rates. There is a critical need for a scalable, genotype-only method that can resolve these "half-degree" ambiguities without requiring observed pedigrees or extensive relative information. ResultsWe present a novel computational framework that achieves near-complete separation of H-S and N/N-A pairs using haplotype-level sharing features derived from across-chromosome phasing. By modeling these features with a multivariate Gaussian mixture model (GMM), we demonstrate exceptional classification performance in biobank-scale data. Ground-truth validation labels were established through a multi-step inference process within a family graph constructed from high-confidence first-degree relationships. To overcome the scarcity of genotyped half-sibling parents, we structurally validated H-S pairs by identifying specific pedigree configurations involving first cousins that are logically incompatible with an avuncular relationship. Our method achieves a sensitivity of 96.9% and a specificity of 99.7% on these structurally-validated pairs. Furthermore, we show that these identified relationships serve as superior phase anchors, that measurably improve the accuracy of across-chromosome homologue assignment. This method provides a robust, scalable solution for pedigree reconstruction and the control of cryptic relatedness in large-scale genomic studies. Contactemmanuel.sapin@colorado.edu, kristen.kelly@colorado.edu, and matthew.c.keller@colorado.edu

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.