DrugSet: A validated R Shiny application for reproducible drug codelist construction from ATC classification to CPRD Aurum prodcodes
Hoxhaj, V.; Fry, C.; Morris, D.; Aurelius, T.; Martin, S.; Sturkenboom, M.; Andaur Navarro, C.
Show abstract
Objectives. To present DrugSet, a validated R Shiny application supporting the construction medicinal products codelists based on the Anatomical Therapeutic Chemical (ATC) system and their mapping to Clinical Practice Research Datalink (CPRD) Aurum prodcodes within a single interactive workflow. Materials and Methods. DrugSet comprises four modules: data preparation, ATC-based hierarchical code selection, string-based CPRD Aurum prodcodes mapping, and codelist export. Validation was conducted against World Health Organization (WHO) ATC reference codelists and manually curated prodcodes mappings across three drug classes: metformin, beta-blocking agents, and topical salicylic acid. Sensitivity, specificity, and Positive Predictive Values (PPV) were calculated for ATC codelist generation. Agreement proportions (overlapping against total identified codes) were calculated for prodcodes mapping. Time needed for codelist construction using DrugSet was recorded and compared to manual approaches. Results. DrugSet ATC codelist generation against WHO manual reference achieved 100% sensitivity, specificity, and PPV across all medicinal products. Prodcodes mapping agreement ranged from 89.2% to 98.3% with discrepancies due to missing data in the prodcodes input vocabulary. DrugSet completed codelist construction in 9 minutes compared to 3 hours and 10 minutes manually, across all medicinal products classes. Discussion. DrugSet provides a unified workflow that runs directly on ATC and source CPRD Aurum vocabulary files. The reduction in codelist construction time and export of the generated codelists supports reproducibility in pharmacoepidemiologic studies where codelist creation can represent a significant proportion of study setup time. Conclusion. DrugSet is an open-source, validated tool that improves accuracy, and efficiency of codelist construction for medicinal products based on ATC codes towards CPRD Aurum prodcodes.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Standardization of drug names in the FDA Adverse Event Reporting System: The DiAna dictionary 97%
- Large-scale empirical identification of candidate comparators for pharmacoepidemiological studies 95%
- Health Information Exchanges (HIEs) as Novel Sources for Population-Based Post Marketing Surveillance of Medical Products: A Pilot Study from the FDA Sentinel Innovation Center 91%
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 94%
- The Medicines Intelligence Data Platform: A population-based data resource from New South Wales, Australia 88%
- Pregnancy pharmacoepidemiology: How often are key methodological elements reported in publications? 88%
Similar papers in this journal
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 92%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 91%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 91%
Similar papers in this journal
Similar papers in this journal
- DrugWAS: Leveraging drug-wide association studies to facilitate drug repurposing for COVID-19 93%
- MTXPK.org: A clinical decision support tool evaluating high-dose methotrexate pharmacokinetics to inform post-infusion care and use of glucarpidase 91%
- Algorithmic identification of treatment-emergent adverse events from clinical notes using large language models: a pilot study in inflammatory bowel disease 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.