Differential Privacy Protection Against Membership Inference Attack on Machine Learning for Genomic Data
Chen, J.; Wang, W. H.; Shi, X.
Show abstract
Machine learning is powerful to model massive genomic data while genome privacy is a growing concern. Studies have shown that not only the raw data but also the trained model can potentially infringe genome privacy. An example is the membership inference attack (MIA), by which the adversary, who only queries a given target model without knowing its internal parameters, can determine whether a specific record was included in the training dataset of the target model. Differential privacy (DP) has been used to defend against MIA with rigorous privacy guarantee. In this paper, we investigate the vulnerability of machine learning against MIA on genomic data, and evaluate the effectiveness of using DP as a defense mechanism. We consider two widely-used machine learning models, namely Lasso and convolutional neural network (CNN), as the target model. We study the trade-off between the defense power against MIA and the prediction accuracy of the target model under various privacy settings of DP. Our results show that the relationship between the privacy budget and target model accuracy can be modeled as a log-like curve, thus a smaller privacy budget provides stronger privacy guarantee with the cost of losing more model accuracy. We also investigate the effect of model sparsity on model vulnerability against MIA. Our results demonstrate that in addition to prevent overfitting, model sparsity can work together with DP to significantly mitigate the risk of MIA.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Privacy-Preserving Federated Neural Network Learning for Disease-Associated Cell Classification 97%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 94%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 93%
Similar papers in this journal
- The Effect of Kinship in Re-identification Attacks Against Genomic Data Sharing Beacons 94%
- Privacy-Preserving and Robust Watermarking on Sequential Genome Data using Belief Propagation and Local Differential Privacy 94%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.