Improve data management in register-based research: Transition from CSV to Parquet
Fenk, S. R.; Furu, K.; Bakken, I. J. L.
Show abstract
AimsTo identify an efficient file format for data delivery from large administrative registers for research, testing the full workflow from data extraction to research data management. MethodsThrough collaboration between a data delivery and a research department within the same institute, we evaluated each step from data extraction and delivery to management and usage, comparing CSV and Parquet formats. ResultsSwitching from the unstructured, text-based CSV format to the highly structured Parquet format significantly optimized all processes by reducing file sizes and saving processing time. The Parquet format also provided access to advanced data management techniques, simplifying further work. Despite these advantages, the basic programming required for Parquet format is not very different from that for CSV. We provide a tutorial and examples as online supplement. ConclusionsWe strongly recommend replacing CSV files with contemporary data formats. The Parquet file format proved to be an excellent option throughout the entire process from data extraction to research implementation
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 97%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 92%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 91%
Similar papers in this journal
- Protocol for Development of a Reporting Guideline for Causal and Counterfactual Prediction Models 92%
- Data-driven discovery of changes in clinical code usage over time: a case-study on changes in cardiovascular disease recording in two English electronic health records databases (2001-2015) 92%
- Cohort Profile: Born in Wales - a birth cohort with maternity, parental, and child data linkage for life course research in Wales, UK 91%
Similar papers in this journal
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 92%
- Scalable information extraction from free text electronic health records using large language models 92%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 91%
Similar papers in this journal
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 93%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.