Back

DeepKin: precise estimation of in-depth relatedness and its application in UK Biobank

Zhang, Q.-X.; Jayasinghe, D.; Lee, S. H.; Xu, H.; Chen, G.-B.

2024-05-01 genetics
10.1101/2024.04.30.591647 bioRxiv
Show abstract

Accurately estimating relatedness between samples is crucial in genetics and epidemiological analysis. Using genome-wide single nucleotide polymorphisms (SNPs), it is now feasible to measure realized relatedness even in the absence of pedigree. However, the sampling variation in SNP-based measures and factors affecting method-of-moments relatedness estimators have not been fully explored, whilst static cut-off thresholds have traditionally been employed to classify relatedness levels for decades. Here, we introduce the deepKin framework as a moment-based relatedness estimation and inference method that incorporates data-specific cut-off threshold determination. It addresses the limitations of previous moment estimators by leveraging the sampling variance of the estimator to provide statistical inference and classification. Key principles in relatedness estimation and inference are provided, including inferring the critical value required to reject the hypothesis of unrelatedness, which we refer to as the deepest significant relatedness, determining the minimum effective number of markers, and understanding the impact on statistical power. Through simulations, we demonstrate that deepKin accurately infers both unrelated pairs and relatives with the support of sampling variance. We then apply deepKin to two subsets of the UK Biobank dataset. In the 3K Oxford subset, tested with four sets of SNPs, the SNP set with the largest effective number of markers and correspondingly the smallest expected sampling variance exhibits the most powerful inference for distant relatives. In the 430K British White subset, deepKin identifies 212,120 pairs of significant relatives and classifies them into six degrees. Additionally, cross-cohort significant relative ratios among 19 assessment centers located in different cities are geographically correlated, while within-cohort analyses indicate both an increase in close relatedness and a potential increase in diversity from north to south throughout the UK. Overall, deepKin presents a novel framework for accurate relatedness estimation and inference in biobank-scale datasets. For biobank-scale application we have implemented deepKin as an R package, available in the GitHub repository (https://github.com/qixininin/deepKin).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Heredity
64 papers in training set
Top 0.1%
14.7%
2
PLOS Genetics
862 papers in training set
Top 0.8%
11.6%
3
Molecular Ecology Resources
171 papers in training set
Top 0.2%
11.6%
4
GENETICS
483 papers in training set
Top 0.5%
10.7%
5
The American Journal of Human Genetics
234 papers in training set
Top 0.7%
6.6%
50% of probability mass above
6
European Journal of Human Genetics
58 papers in training set
Top 0.2%
5.3%
7
Frontiers in Genetics
230 papers in training set
Top 0.9%
3.9%
8
G3 Genes|Genomes|Genetics
351 papers in training set
Top 1%
3.3%
9
Genetic Epidemiology
55 papers in training set
Top 0.2%
3.1%
10
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
11
Behavior Genetics
17 papers in training set
Top 0.1%
2.4%
12
Molecular Biology and Evolution
542 papers in training set
Top 3%
2.1%
13
G3: Genes, Genomes, Genetics
252 papers in training set
Top 2%
1.9%
14
eLife
5828 papers in training set
Top 50%
1.7%
15
Bioinformatics
1204 papers in training set
Top 7%
1.3%
16
Nature Communications
5641 papers in training set
Top 49%
1.3%
17
PLOS Computational Biology
1863 papers in training set
Top 18%
1.1%
18
Human Molecular Genetics
141 papers in training set
Top 2%
1.1%
19
Genetics Selection Evolution
39 papers in training set
Top 0.2%
1.1%
20
Genome Research
468 papers in training set
Top 5%
1.1%
21
Human Genetics
28 papers in training set
Top 0.5%
1.0%
22
Scientific Reports
3612 papers in training set
Top 76%
0.8%
23
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
24
Human Genetics and Genomics Advances
84 papers in training set
Top 2%
0.6%
25
Molecular Ecology
336 papers in training set
Top 4%
0.6%
26
Genes
144 papers in training set
Top 5%
0.6%