Back

OpenCodeCounts: An open-access, interactive online tool and R package for analysing clinical code usage in England

Tamborska, A. A.; Higgins, R.; Boukari, Y.; Kingsley, V.; Ojedele, L.; Oreagba, K.; Massey, J.; Schaffer, A.; Green, A.; Hulme, W.; MacKenna, B.; Curtis, H. J.; Fisher, L.; Wiedemann, M.

2025-10-15 health informatics
10.1101/2025.10.14.25338005 medRxiv
Show abstract

Clinical codes are unique identifiers used in electronic health records to document specific information, such as diagnoses, procedures or medications. Because of their structured and systematic nature, they are often used for research, audit and service evaluation. Knowing how frequently certain codes are recorded can be invaluable in planning such work. For example, not all events are recorded with equal frequency, so knowing the usage of specific codes helps determine research feasibility. In England, data on the frequency of clinical code recording is openly available for three classification systems used in primary and secondary care: SNOMED CT, ICD-10 and OPCS-4. However, these valuable datasets, showing how frequently individual clinical code were used, are difficult to access and analyse in their current format, hindering their application in research and service evaluation. We developed OpenCodeCounts, an interactive online tool with an accompanying R package that help users explore and analyse Englands primary and secondary care clinical coding data. Both are compatible with OpenCodelists.org: a publicly available and free-to-use platform for codelist development and sharing. This article describes the underlying datasets, the development of the opencodecounts R package and the interactive app, and showcases their applications for electronic health records research. Plain English SummaryElectronic health records (EHR) are a summary of patients medical files. They are a useful resource for researchers who want to study health and healthcare use. EHR often use unique codes to record diagnoses, treatments, procedures and other clinical information. These codes help organise and standardise healthcare data. Researchers working with EHR need to know how often different codes are used to plan studies and build accurate lists of codes (called codelists) for their research. In England, data on how often these medical codes are used is publicly available as large spreadsheets, which makes it time-consuming to access and analyse. To make this easier, we created an interactive online tool that allows researchers to explore and visualise this information easily, as well as an R package - a collection of tools written in the R programming language. The package, called opencodecounts, allows researchers to access and study this information. We made these tools compatible with a publicly available website for codelists preparation and sharing, called OpenCodelists.org. In this article, we describe the underlying data, explain how the tools were developed, and demonstrate how they can be used to support research using EHR.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.