Genomic privacy risks in GWAS summary statistics
Lan, A.; Pawitan, Y.; Shen, X.
Show abstract
The rapid advancement in sequencing technologies has exponentially increased the availability of genomic data, heightening concerns about data privacy. Despite the perceived safety of publicly accessible genome-wide association study (GWAS) summary statistics, we demonstrate that their combination with less sensitive high-dimensional phenotype data can lead to significant leakage of confidential genomic information. By transforming a linear regression model into linear programming constraints, we scrutinize the potential for genomic data recovery using GWAS summary statistics. We found that an effective phenotype-to-sample size ratio above 0.85 could enable full genotype recovery, and that above 0.16 was sufficient to enable individual identification. Certain non-European populations are especially vulnerable. The results stress the urgent need for stronger privacy protections in genomic research while maintaining data utility.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast estimation of genetic correlation for Biobank-scale data 96%
- Probabilistic Colocalization of Genetic Variants from Complex and Molecular Traits: Promise and Limitations 96%
- Analyzing and Reconciling Colocalization and Transcriptome-wide Association Studies from the Perspective of Inferential Reproducibility 96%
Similar papers in this journal
- Integrating Comprehensive Functional Annotations to Boost Power and Accuracy in Gene-Based Association Analysis 97%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
- Leveraging expression from multiple tissues using sparse canonical correlation analysis (sCCA) and aggregate tests improves the power of transcriptome-wide association studies (TWAS) 96%
Similar papers in this journal
- Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores 96%
- A parametric bootstrap approach for computing confidence intervals for genetic correlations with application to genetically-determined protein-protein networks 94%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.