OpenCodeCounts: An open-access, interactive online tool and R package for analysing clinical code usage in England
Tamborska, A. A.; Higgins, R.; Boukari, Y.; Kingsley, V.; Ojedele, L.; Oreagba, K.; Massey, J.; Schaffer, A.; Green, A.; Hulme, W.; MacKenna, B.; Curtis, H. J.; Fisher, L.; Wiedemann, M.
Show abstract
Clinical codes are unique identifiers used in electronic health records to document specific information, such as diagnoses, procedures or medications. Because of their structured and systematic nature, they are often used for research, audit and service evaluation. Knowing how frequently certain codes are recorded can be invaluable in planning such work. For example, not all events are recorded with equal frequency, so knowing the usage of specific codes helps determine research feasibility. In England, data on the frequency of clinical code recording is openly available for three classification systems used in primary and secondary care: SNOMED CT, ICD-10 and OPCS-4. However, these valuable datasets, showing how frequently individual clinical code were used, are difficult to access and analyse in their current format, hindering their application in research and service evaluation. We developed OpenCodeCounts, an interactive online tool with an accompanying R package that help users explore and analyse Englands primary and secondary care clinical coding data. Both are compatible with OpenCodelists.org: a publicly available and free-to-use platform for codelist development and sharing. This article describes the underlying datasets, the development of the opencodecounts R package and the interactive app, and showcases their applications for electronic health records research. Plain English SummaryElectronic health records (EHR) are a summary of patients medical files. They are a useful resource for researchers who want to study health and healthcare use. EHR often use unique codes to record diagnoses, treatments, procedures and other clinical information. These codes help organise and standardise healthcare data. Researchers working with EHR need to know how often different codes are used to plan studies and build accurate lists of codes (called codelists) for their research. In England, data on how often these medical codes are used is publicly available as large spreadsheets, which makes it time-consuming to access and analyse. To make this easier, we created an interactive online tool that allows researchers to explore and visualise this information easily, as well as an R package - a collection of tools written in the R programming language. The package, called opencodecounts, allows researchers to access and study this information. We made these tools compatible with a publicly available website for codelists preparation and sharing, called OpenCodelists.org. In this article, we describe the underlying data, explain how the tools were developed, and demonstrate how they can be used to support research using EHR.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 93%
- Using data science to identify unusual treatment choices in England: illustrative findings of uncommon antipsychotics Pericyazine and Promazine 93%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 93%
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
- Assessing the quality of clinical and administrative data extracted from hospitals: The General Medicine Inpatient Initiative (GEMINI) experience 93%
- medExtractR: A medication extraction algorithm for electronic health records using the R programming language 93%
Similar papers in this journal
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 93%
- Development of a data-driven COVID-19 prognostication tool to inform triage and step-down care for hospitalised patients in Hong Kong: A population based cohort study 92%
- A Multi-Granular Stacked Regression for Forecasting Long-Term Demand in Emergency Departments 91%
Similar papers in this journal
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 93%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 92%
- LinkR: an open source, low-code and collaborative data science platform for healthcare data analysis and visualization 92%
Similar papers in this journal
- Data-driven discovery of changes in clinical code usage over time: a case-study on changes in cardiovascular disease recording in two English electronic health records databases (2001-2015) 94%
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 92%
- Ethnic variation in outcome of people hospitalised with Covid-19 in Wales (UK): A rapid analysis of surveillance data using Onomap, a name-based ethnicity classification tool 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.