Back

CodeMergeR: A shiny-based application for codelist integration

Hoxhaj, V.; Riera Arnau, J.; Mohammadi, S.; Sturkenboom, M. C.; Andaur Navarro, C. L.

2025-12-08 epidemiology
10.64898/2025.12.07.25341784 medRxiv
Show abstract

ObjectiveTo describe CodeMergeR, an open-source R shiny application for the standardization, cleaning, and integration of individual concept codelists into a master file for real-world evidence (RWE) studies. Material and MethodsCodeMergeR was designed to align with concept codelists that are outputted by CodeMapper v1.0. The application includes two main modules: (1) Conformance and Coherence, which validates file naming conventions, metadata consistency, and structural integrity; and (2) Cleaning and Standardization, which removes duplicates, corrects formatting issues, and merges validated codelists. Functional and performance testing were conducted using real-world library metadata and codelists obtained from the VAC4EU SharePoint library, including manually introduced errors to assess detection capabilities. ResultsThe CodeMergeR application successfully identified and resolved a wide range of structural and semantic issues, achieving an overall error detection rate of 96.8% out of 38 errors in the functional testing. The application processed over 50 clinical concept folders and 10 algorithms in under one minute during performance testing. Key issues detected were related to file names, duplications of concepts, formatting anomalies of codes (e.g., scientific notation, rounding), and concept metadata mismatches. DiscussionCodeMergeR shows the need for reproducible and scalable concept codelist management for RWE. In contrast to existing tools such as CodelistGenerator, ATHENA, and OpenCodelists, CodeMergeR fills an important quality aspect in RWE generation and transparency. ConclusionCodeMergeR contributes to improved reproducibility, scalability, and transparency in individual codelist management in RWE studies. KEY POINTSO_LIHealth data collected via electronic records (real-world data, RWD) have shown to be a valuable source of information to generate real world evidence (RWE) on effectiveness and safety of medicines and vaccines. C_LIO_LIVery little guidance exists on codelist integration, even though they are essential for semantic harmonization and identification of clinical concepts in RWD. Their manual integration is error-prone and time-consuming, making this an important and under-addressed opportunity to improve RWE quality. C_LIO_LICodeMergeR is an open-source R Shiny application developed to standardize, verify, and integrate individual concept codelists into a harmonized master file for analytical pipelines in RWE studies. C_LI PLAIN LANGUAGE SUMMARYResearchers who study the real-world use and effects of medicines often use data from electronic health records. This type of data, called real-world data (RWD), is generally analyzed in a federated manner (data stay local) using a common protocol, common analytics and a common data model. In this way RWD from different geographical regions and systems can be used even if diagnoses and treatments are recorded in different vocabularies. To make sure these can be retrieved, compared and combined, researchers use codelists, a set of files with a list of medical codes related to the clinical event of interest. However, concatenation of files can be difficult and time-consuming due to differences in formats, naming styles over time and several people contributing to its creation. To help with this, we created CodeMergeR, a free, open-source tool that helps researchers clean, check, and combine codelists automatically. It finds and fixes common problems, like missing files, different naming styles, or formatting errors. CodeMergeR fits easily into the tools built on the ConcePTION common data model by the Vaccine Monitoring Collaboration for Europe (VAC4EU) and has been shown to detect almost all issues (96.8%) quickly and accurately. CodeMergeR helps make studies using codelists more reliable, faster, and easier to repeat. In the future, it could also be used in other research networks that use multiple vocabularies.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.