Automated Extraction and Classification of Drug Prescriptions in Electronic Health Records: Introducing the PRESNER Pipeline
Colon-Ruiz, C.; Fitzgerald, T. W.; Segura-Bedmar, I.; Birney, E.; Herrero-Zazo, M.
Show abstract
Electronic health record (EHR) systems with prescription data offer vast potential in pharmacoepidemiology and pharmacogenomics. The large amount of clinical data recorded in these systems requires automatic processing to extract relevant information. This paper introduces PRESNER, a name entity recognition (NER) and classification pipeline for EHR prescription data. The pipeline uses the pre-trained transformer Bio-ClinicalBERT fine-tuned on UK Biobank prescription entries manually annotated with medication-related information (drug name, route of administration, pharmaceutical form, strength, and dosage) as the core NER system. Moreover, PRESNER also maps drugs to the Anatomical Therapeutic and Chemical (ATC) classification system and distinguishes between systemic and non-systemic drug products. It outperformed a baseline model combining the state-of-the-art Med7 and a dictionary-based approach from the ChEMBL database with a macro-average F1-score of 0.95 vs 0.71. In addition to UK Biobank prescription data, PRESNER can also be applied to other English prescription datasets, making it a versatile tool for researchers in the field.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Active Neural Networks to Detect Mentions of Changes to Medication Treatment in Social Media 93%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 92%
- medExtractR: A medication extraction algorithm for electronic health records using the R programming language 92%
Similar papers in this journal
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 89%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 89%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 92%
- Using indication embeddings to represent patient health for drug safety studies 92%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.