Data Anonymization for Open Science: A Case Study
Francis, P.; Jurak, G.; Leskosek, B.; Prasser, F.; Otte, K.
Show abstract
One of many challenges to open science is anonymization of personal data so that it may be shared. This paper presents a case study of the anonymization of a dataset containing cardio-respiratory fitness and commuting patterns for Slovenian school children. It evaluates three different anonymization tools, ARX, SDV, and SynDiffix. The fitness study was selected because its small size (N=713) and generally low statistical significance make it particularly challenging for data anonymization. Unlike most prior anonymization tool evaluations, this paper examines whether the scientific conclusions of the original study would have been supported by the anonymized datasets. It also considers the burden imposed on researchers using the tools both for data generation and data analysis.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 96%
- Academic Tracker: Software for Tracking and Reporting Publications Associated with Authors and Grants 94%
- Understanding signaling and metabolic paths using semantified and harmonized information about biological interactions 94%
Similar papers in this journal
Similar papers in this journal
- Emulation of epidemics via Bluetooth-based virtual safe virus spread: experimental setup, software, and data 94%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
- A data management system for precision medicine 91%
Similar papers in this journal
- COVID19-Global: A shiny application to perform a global comparative data visualization for the SARS-CoV-2 epidemic 93%
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 93%
- Data-Driven Prediction of COVID-19 Cases in Germany for Decision Making 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.