Back

Evaluating anonymized genome re-identification using polygenic predictions and its implications for data privacy

Cavinato, T.; Hofmeister, R. J.; Kutalik, Z.

2026-06-10 genetics
10.64898/2026.06.10.731306 bioRxiv
Show abstract

Re-identification by phenotypic prediction aims to determine whether a genome belongs to a specific individual by comparing the individuals known traits with those predicted from the genome. This type of tracing attack is widely discussed in the genomic privacy literature, yet previous studies have been criticized for overstating its practical risks. Over the past decade, genome-wide association studies (GWAS) with increasing sample size improved the accuracy of phenotypic prediction, potentially enhancing such attacks. To quantify their real-world threat, we developed a probabilistic framework that estimates the likelihood of a match between an individuals observed traits and polygenic scores (PGS) derived from a genome, while accounting for prediction accuracy and genetic and environmental correlations between the traits. We benchmarked this re-identification method and examined how the prior probability (reflecting the a priori chance that a random genome and set of traits correspond to the same person) affects performance. Finally, we assessed whether sensitive information could be inferred through this attack by attempting to predict multiple sensitive haplotypes, such as APOE-{varepsilon}4 (linked with Alzheimers disease). Our re-identification method outperformed a state-of-the-art tool, and reached a precision above 99% for a recall of 40% when considering a prior of 50%. However, after considering real-world settings, we estimated that realistic priors would not exceed 4 x 10-4%, resulting in a precision lower than 0.13% at the same recall (40%). The inference of sensitive genotypes also proved ineffective, as achieving a precision above 50% for identifying APOE-{varepsilon}4 carriers was only possible at a recall below 20%. To conclude, although re-identification by phenotypic prediction is technically feasible, our findings indicate that its effectiveness in real-world conditions is limited. These results counterpoint to earlier claims of severe genomic privacy risks and offer guidance for policymakers, biobank administrators, and research participants.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Forensic Science International: Genetics
26 papers in training set
Top 0.1%
11.9%
2
Frontiers in Genetics
230 papers in training set
Top 0.1%
11.9%
3
Nature Communications
5641 papers in training set
Top 24%
6.7%
4
Scientific Reports
3612 papers in training set
Top 13%
6.2%
5
Patterns
78 papers in training set
Top 0.2%
5.5%
6
The American Journal of Human Genetics
234 papers in training set
Top 0.8%
5.5%
7
European Journal of Human Genetics
58 papers in training set
Top 0.2%
5.5%
50% of probability mass above
8
Bioinformatics
1204 papers in training set
Top 5%
4.0%
9
PLOS ONE
5266 papers in training set
Top 42%
2.4%
10
Genome Research
468 papers in training set
Top 3%
2.1%
11
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 24%
2.1%
12
Briefings in Bioinformatics
354 papers in training set
Top 4%
1.7%
13
Human Genetics and Genomics Advances
84 papers in training set
Top 1%
1.7%
14
Cell Genomics
172 papers in training set
Top 2%
1.7%
15
iScience
1154 papers in training set
Top 17%
1.7%
16
Genome Biology
637 papers in training set
Top 6%
1.5%
17
Science Advances
1243 papers in training set
Top 22%
1.4%
18
GigaScience
212 papers in training set
Top 3%
1.4%
19
Cell
431 papers in training set
Top 7%
1.3%
20
eLife
5828 papers in training set
Top 57%
1.1%
21
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.1%
22
Nucleic Acids Research
1281 papers in training set
Top 12%
1.0%
23
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
24
BMC Genomics
406 papers in training set
Top 8%
0.8%
25
Nature Genetics
286 papers in training set
Top 5%
0.8%
26
npj Digital Medicine
118 papers in training set
Top 3%
0.8%
27
Genetic Epidemiology
55 papers in training set
Top 0.8%
0.6%
28
Nature Methods
385 papers in training set
Top 7%
0.6%
29
PLOS Genetics
862 papers in training set
Top 13%
0.6%
30
PLOS Biology
486 papers in training set
Top 14%
0.6%