Development and validation of MedDRA Tagger: a tool for extraction and structuring medical information from clinical notes
Humbert-Droz, M.; Corley, J.; Tamang, S.; Gevaert, O.
Show abstract
Rapid and automated extraction of clinical information from patients notes is a desirable though difficult task. Natural language processing (NLP) and machine learning have great potential to automate and accelerate such applications, but developing such models can require a large amount of labeled clinical text, which can be a slow and laborious process. To address this gap, we propose the MedDRA tagger, a fast annotation tool that makes use of industrial level libraries such as spaCy, biomedical ontologies and weak supervision to annotate and extract clinical concepts at scale. The tool can be used to annotate clinical text and obtain labels for training machine learning models and further refine the clinical concept extraction performance, or to extract clinical concepts for observational study purposes. To demonstrate the usability and versatility of our tool, we present three different use cases: we use the tagger to determine patients with a primary brain cancer diagnosis, we show evidence of rising mental health symptoms at the population level and our last use case shows the evolution of COVID-19 symptomatology throughout three waves between February 2020 and October 2021. The validation of our tool showed good performance on both specific annotations from our development set (F1 score 0.81) and open source annotated data set (F1 score 0.79). We successfully demonstrate the versatility of our pipeline with three different use cases. Finally, we note that the modular nature of our tool allows for a straightforward adaptation to another biomedical ontology. We also show that our tool is independent of EHR system, and as such generalizable.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 96%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 95%
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 93%
Similar papers in this journal
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
- PheMIME: An Interactive Web App and Knowledge Base for Phenome-Wide, Multi-Institutional Multimorbidity Analysis 93%
Similar papers in this journal
Similar papers in this journal
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 93%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 93%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.