PRISM-G: an interpretable privacy scoring method for assessing risk in synthetic human genome data
Correa Rojo, A.; Moreau, Y.; Ertaylan, G.
Show abstract
The growing use of synthetic genomic data promises broader data access but raises unresolved concerns about privacy risk. We introduce PRISM-G, a model-agnostic framework that summarizes privacy exposure of synthetic genomes across three complementary components: (i) a proximity view that asks whether synthetic individuals lie unusually close to real genomes in genetic-coordinate space; (ii) a kinship view that detects replay of familial or population-structure patterns beyond what is expected by chance; and (iii) a trait-linked view that captures exposure through rare variants and simple membership-inference signals. Each component yields a normalized risk score and a risk-averse aggregation maps these to a 0-100 PRISM-G score. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT-solver (Genomator). Our results show that privacy vulnerabilities concentrate along different axes across models and marker densities, underscoring that a single privacy-based similarity metric is insufficient.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Assessing transcriptomic re-identification risks using discriminative sequence models 96%
- Secure Discovery of Genetic Relatives across Large-Scale and Distributed Genomic Datasets 96%
- ML-MAGES: A machine learning framework for multivariate genetic association analyses with genes and effect size shrinkage 95%
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 95%
- Causal differential expression analysis under unmeasured confounders with causarray 93%
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.