Methods for safely sharing dual-use genetic data
Sawaya, S.; Lo, C.-C.; Li, P.-E.; Hovde, B.; Chain, P.
Show abstract
AbstractO_ST_ABSBackgroundC_ST_ABSSome genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm. ResultsHere we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads. ConclusionOur methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 95%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 94%
Similar papers in this journal
Similar papers in this journal
- Scaling Logical Density of DNA storage with Enzymatically-Ligated Composite Motifs 93%
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 92%
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.