CodeMergeR: A shiny-based application for codelist integration
Hoxhaj, V.; Riera Arnau, J.; Mohammadi, S.; Sturkenboom, M. C.; Andaur Navarro, C. L.
Show abstract
ObjectiveTo describe CodeMergeR, an open-source R shiny application for the standardization, cleaning, and integration of individual concept codelists into a master file for real-world evidence (RWE) studies. Material and MethodsCodeMergeR was designed to align with concept codelists that are outputted by CodeMapper v1.0. The application includes two main modules: (1) Conformance and Coherence, which validates file naming conventions, metadata consistency, and structural integrity; and (2) Cleaning and Standardization, which removes duplicates, corrects formatting issues, and merges validated codelists. Functional and performance testing were conducted using real-world library metadata and codelists obtained from the VAC4EU SharePoint library, including manually introduced errors to assess detection capabilities. ResultsThe CodeMergeR application successfully identified and resolved a wide range of structural and semantic issues, achieving an overall error detection rate of 96.8% out of 38 errors in the functional testing. The application processed over 50 clinical concept folders and 10 algorithms in under one minute during performance testing. Key issues detected were related to file names, duplications of concepts, formatting anomalies of codes (e.g., scientific notation, rounding), and concept metadata mismatches. DiscussionCodeMergeR shows the need for reproducible and scalable concept codelist management for RWE. In contrast to existing tools such as CodelistGenerator, ATHENA, and OpenCodelists, CodeMergeR fills an important quality aspect in RWE generation and transparency. ConclusionCodeMergeR contributes to improved reproducibility, scalability, and transparency in individual codelist management in RWE studies. KEY POINTSO_LIHealth data collected via electronic records (real-world data, RWD) have shown to be a valuable source of information to generate real world evidence (RWE) on effectiveness and safety of medicines and vaccines. C_LIO_LIVery little guidance exists on codelist integration, even though they are essential for semantic harmonization and identification of clinical concepts in RWD. Their manual integration is error-prone and time-consuming, making this an important and under-addressed opportunity to improve RWE quality. C_LIO_LICodeMergeR is an open-source R Shiny application developed to standardize, verify, and integrate individual concept codelists into a harmonized master file for analytical pipelines in RWE studies. C_LI PLAIN LANGUAGE SUMMARYResearchers who study the real-world use and effects of medicines often use data from electronic health records. This type of data, called real-world data (RWD), is generally analyzed in a federated manner (data stay local) using a common protocol, common analytics and a common data model. In this way RWD from different geographical regions and systems can be used even if diagnoses and treatments are recorded in different vocabularies. To make sure these can be retrieved, compared and combined, researchers use codelists, a set of files with a list of medical codes related to the clinical event of interest. However, concatenation of files can be difficult and time-consuming due to differences in formats, naming styles over time and several people contributing to its creation. To help with this, we created CodeMergeR, a free, open-source tool that helps researchers clean, check, and combine codelists automatically. It finds and fixes common problems, like missing files, different naming styles, or formatting errors. CodeMergeR fits easily into the tools built on the ConcePTION common data model by the Vaccine Monitoring Collaboration for Europe (VAC4EU) and has been shown to detect almost all issues (96.8%) quickly and accurately. CodeMergeR helps make studies using codelists more reliable, faster, and easier to repeat. In the future, it could also be used in other research networks that use multiple vocabularies.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 96%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 91%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 91%
Similar papers in this journal
Similar papers in this journal
- Scalable information extraction from free text electronic health records using large language models 93%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 93%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 91%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.