dsSurvival 2.0: Privacy enhancing survival curves for survival models in the federated DataSHIELD analysis system
Banerjee, S.; Bishop, T.
Show abstract
ObjectiveSurvival models are used extensively in biomedical sciences, where they allow the investigation of the effect of exposures on health outcomes. It is desirable to use diverse data sets in survival analyses, because this offers increased statistical power and generalisability of results. However, there are often challenges with bringing data together in one location or following an analysis plan and sharing results. DataSHIELD is an analysis platform that helps users to overcome these ethical, governance and process difficulties. It allows users to analyse data remotely, using functions that are built to restrict access to the detailed data items (federated analysis). Previous works have provided survival modelling functionality in DataSHIELD (dsSurvival package), but there is a requirement to provide functions that offer privacy enhancing survival curves that retain useful information. ResultsWe introduce an enhanced version of the dsSurvival package which offers privacy enhancing survival curves for DataSHIELD. Different methods for enhancing privacy were evaluated for their effec-tiveness in enhancing privacy while maintaining utility. We demonstrated how our selected method could enhance privacy in different scenarios using real survival data. The details of how DataSHIELD can be used to generate survival curves can be found in the associated tutorial.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accessibility of covariance information creates vulnerability in Federated Learning frameworks 95%
- Privacy-Preserving and Robust Watermarking on Sequential Genome Data using Belief Propagation and Local Differential Privacy 95%
- The Effect of Kinship in Re-identification Attacks Against Genomic Data Sharing Beacons 93%
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 92%
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 92%
- Advancing clinical cohort selection with genomics analysis on a distributed platform 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.