Back

Improve data management in register-based research: Transition from CSV to Parquet

Fenk, S. R.; Furu, K.; Bakken, I. J. L.

2025-10-17 epidemiology
10.1101/2025.10.15.25337992 medRxiv
Show abstract

AimsTo identify an efficient file format for data delivery from large administrative registers for research, testing the full workflow from data extraction to research data management. MethodsThrough collaboration between a data delivery and a research department within the same institute, we evaluated each step from data extraction and delivery to management and usage, comparing CSV and Parquet formats. ResultsSwitching from the unstructured, text-based CSV format to the highly structured Parquet format significantly optimized all processes by reducing file sizes and saving processing time. The Parquet format also provided access to advanced data management techniques, simplifying further work. Despite these advantages, the basic programming required for Parquet format is not very different from that for CSV. We provide a tutorial and examples as online supplement. ConclusionsWe strongly recommend replacing CSV files with contemporary data formats. The Parquet file format proved to be an excellent option throughout the entire process from data extraction to research implementation

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.