Varlock: privacy preserving storage and dissemination of sequenced genomic data
Hekel, R.; Budis, J.; Kucharik, M.; Radvanszky, J.; Szemes, T.
Show abstract
IntroductionCurrent and future applications of genomic data may raise ethical and privacy concerns. Processing and storing these data introduces a risk of abuse by a potential adversary since a human genome contains sensitive personal information. For this reason, we developed a privacy preserving method, called Varlock, for secure storage of sequenced genomic data. Materials and methodsWe used a public set of population allele frequencies to mask personal alleles detected in genomic reads. Each personal allele described by the public set is masked by a randomly selected population allele with respect to its frequency. Masked alleles are preserved in an encrypted confidential file that can be shared, in whole or in part, using public-key cryptography. ResultsOur method masked personal variants and introduced new variants detected in a personal masked genome. Alternative alleles with lower population frequency were masked and introduced more often. We performed a joint PCA analysis of personal and masked VCFs, showing that the VCFs between the two groups can not be trivially mapped. Moreover, the method is reversible and personal alleles can be unmasked in specific genomic regions on demand. ConclusionOur method masks personal alleles within genomic reads while preserving valuable non-sensitive properties of sequenced DNA fragments for further research. Personal alleles may be restored in desired genomic regions and shared with patients, clinics, and researchers. We suggest that the method can provide an additional layer of security for storing and sharing the raw aligned reads.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
- T1K: efficient and accurate KIR and HLA genotyping with next-generation sequencing data 95%
- HISAT-3N: a rapid and accurate three-nucleotide sequence aligner 95%
Similar papers in this journal
- Vulcan: Improved long-read mapping and structural variant calling via dual-mode alignment 95%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 95%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.