Differentially Private Distributed Inference
Papachristou, M.; Rahimian, M. A.
Show abstract
How can agents exchange information to learn from each other despite their privacy needs and security concerns? Consider healthcare centers that want to collaborate on a multicenter clinical trial, but are concerned about sharing sensitive patient information. Preserving individual privacy and enabling efficient social learning are both important desiderata, but they seem fundamentally at odds. We attempt to reconcile these desiderata by controlling information leakage using statistical disclosure control methods based on differential privacy (DP). Our agents use log-linear rules to update their belief statistics after communicating with their neighbors. DP randomization of beliefs offers communicating agents with plausible deniability with regard to their private information and is amenable to rigorous performance guarantees for the quality of statistical inference. We consider two information environments: one for distributed maximum likelihood estimation (MLE) given a finite number of private signals available at the start of time and another for online learning from an infinite, intermittent stream of private signals that arrive over time. Noisy information aggregation in the finite case leads to interesting trade-offs between rejecting low-quality states and making sure that all high-quality states are admitted in the algorithm output. The MLE setting has natural applications to binary hypothesis testing that we formalize with relevant statistical guarantees. Our results flesh out the nature of the trade-offs between the quality of the inference, learning accuracy, communication cost, and the level of privacy protection that the agents are afforded. In simulation studies, we perform a differentially private, distributed survival analysis on real-world data from an AIDS Clinical Trials Group (ACTG) study to determine whether new treatments improve over standard care. In addition, we used data from clinical trials in advanced cancer patients to determine whether certain biomedical indices affect patient survival. We show that our methods can achieve privacy-preserving inference with significantly more efficient computations than existing privacy-aware methods based on homomorphic encryption, and at lower error rates compared to first-order differentially private distributed optimization methods.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Expected 10-anonymity of HyperLogLog sketches for federated queries of clinical data repositories 97%
- TiTUS: Sampling and Summarizing Transmission Trees with Multi-strain Infections 94%
- Privacy-Preserving and Robust Watermarking on Sequential Genome Data using Belief Propagation and Local Differential Privacy 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Federated queries of clinical data repositories: balancing accuracy and privacy 95%
- Dynamics and Development of the COVID-19 Epidemics in the US: a Compartmental Model with Deep Learning Enhancement 88%
- Optimizing the Implementation of Clinical Predictive Models to Minimize National Costs: A Sepsis Case Study 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.