Use of commercially available natural language processing software to identify bleeding from the medical record
Walker, A. L.; Watson, C.; Butcher, R.; Butcher, R.; Yandell, M.; Shah, R. U.
Show abstract
BackgroundReal-world evidence derived from the electronic medical record (EMR) is increasingly prevalent. How best to ascertain cardiovascular outcomes from EMRs is unknown. We sought to validate a commercially available natural language processing (NLP) software to extract bleeding events. MethodsWe included patients with atrial fibrillation and cancer seen at our cancer center from 1/1/2016 to 12/31/2019. A query set based on SNOMED CT expressions was created to represent bleeding from 11 different organ systems. We ran the query against the clinical notes and randomly selected a sample of notes for physician validation. The primary outcome was the positive predictive value (PPV) of the software to identify bleeding events stratified by organ system. ResultsWe included 1370 patients with mean age 72 years old (SD 1.5) and 35% female. We processed 66,130 notes; the NLP software identified 6522 notes including 654 unique patients with possible bleeding events. Among 1269 randomly selected notes, the PPV of the software ranged from 0.921 for neurologic bleeds to 0.571 for OB/GYN bleeds. Patterns related to false positive bleeding events identified by the software included historic bleeds, hypothetical bleeds, missed negatives, and word errors. ConclusionsNLP may provide an alternative for population-level screening for bleeding outcomes in cardiovascular studies. Human validation is still needed, but an NLP-driven screening approach may improve efficiency.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 91%
Similar papers in this journal
- OpenSAFELY: impact of national guidance on switching from warfarin to direct oral anticoagulants (DOACs) in early phase of COVID-19 pandemic in England 93%
- The Impact of COVID-19 Pandemic on Cardiology Services 91%
- Anticoagulation trends in adults aged 65 and over with atrial fibrillation; a cohort study 91%
Similar papers in this journal
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 93%
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 92%
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 94%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 94%
- CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.