T-Rx: A toolbox for reproducible processing of prescriptions (Rx) from electronic health records
Lo, C. W. H.; Handley, D.; Pain, O.; Kamp, M.; Gillett, A. C.; Iveson, M. H.; Fabbri, C.; Young, K. G.; AMBER Research Team, ; Lewis, C. M.
Show abstract
Linkage between population-wide biobanks and electronic health records (EHRs) opens new opportunities to study the genetic and epidemiological underpinnings of treatment outcomes, a key step forward in delivering precision medicine. However, challenges include complexities in data extraction for longitudinal analyses, and the absence of reproducible phenotyping algorithms. Here, we present T-Rx, an open-source R package to streamline the processing of prescription and dispensing records, to enable reproducible and scalable analysis across EHR databases. T-Rx consists of three modules to derive treatment-related phenotypes from uncleaned prescription records: (1) Extraction and Imputation Module for extracting and imputing prescription details; (2) Exposure Ascertainment Module for converting prescriptions to longitudinal exposure periods; and (3) Phenotyping Module for creating reproducible proxy phenotypes that capture treatment and response patterns. We tested the utility of T-Rx in UK Biobank primary care records, with strength and quantity information extracted and imputed for 2,721,921 antidepressant and 430,705 antipsychotic prescriptions. The extraction functions were validated using oral hypoglycemic agent prescriptions in Clinical Practice Research Datalink Aurum, showing comparable performances to extraction using the NHS Dictionary of Medicines and Devices (dm+d) codes. The Exposure Ascertainment Module of T-Rx converts discrete prescription or dispensing events into longitudinal exposure periods in one-line R commands, with customizable parameters to account for real-world treatment complexities. The Phenotyping Module takes prescriptions as direct user input and returns analysis -ready data frames. Current phenotyping algorithms include antidepressant switching and treatment-resistant depression. The phenotyping functions also allow flexible parameter choices, such as treatment episode windows, definitions for switching and quality control criteria. Researchers can contribute phenotyping algorithms to T-Rx for reproducible use. T-Rx improves the accessibility of prescription information in biobanks, and analysis of dosage- and treatment-patterns across therapeutic areas. T-Rx contributes to open science through harmonized phenotypic definitions and reproducible analyses of proxy treatment outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- medExtractR: A medication extraction algorithm for electronic health records using the R programming language 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 92%
Similar papers in this journal
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 95%
- Using indication embeddings to represent patient health for drug safety studies 93%
- Framework for Identifying Drug Repurposing Candidates from Observational Healthcare Data 92%
Similar papers in this journal
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 92%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 91%
Similar papers in this journal
- An Informatics Consult approach for generating clinical evidence for treatment decisions 92%
- Explainable AI enables clinical trial patient selection to retrospectively improve treatment effects in schizophrenia 90%
- Combining structured and unstructured data in eMRs to create clinically-defined eMR-derived cohorts 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.