Creating a scalable CT yield metric for pulmonary embolisms in the emergency department using an open-source large language model
Wiseman, B.; Li, H.; Mason, E.; Lonergan, K.; Mitchell, R.; Hayward, J.
Show abstract
BackgroundCT scans are the gold-standard diagnostic test for pulmonary embolisms (PE). Despite stable PE prevalence, CT use is rising in emergency departments (EDs), suggesting test overuse. Current methods for measuring test yield are error-prone or not scalable, thus we tested the accuracy of an open-source, foundational large language model (LLM) for identifying PEs from free-text radiology reports. MethodsOur retrospective diagnostic accuracy study used 10,173 CT-PE reports from 216 radiologists at 38 EDs across Alberta, Canada from April 2021-April 2023. Reports were classified as PE present, PE absent, or Indeterminate by human labelers. An LLM (LLAMA-2-70B) was then prompt-engineered to label the reports. Label accuracy was compared against ICD-10-CA codes and a rule-based natural language processing (NLP) algorithm (ChartExtract; University of Toronto). Descriptive statistics were performed to analyze results. Results1070 (11.8%) reports were PE positive. The LLM achieved an Area Under the Curve (AUC) of 99.1%, outperforming both ICD-10-CA (AUC=90.6%) and ChartExtract (AUC=86.5%), while demonstrating a 16-25% higher sensitivity (LLM: sensitivity=98.8%, specificity=99.1%; ICD-10-CA: sensitivity=82.0%, specificity=99.1%; ChartExtract: sensitivity=73.6%, specificity=99.4%). The LLM took an average of 143 milliseconds to label each report and produce a paragraph justifying the classification. ConclusionsOpen-source, foundational LLMs are an accurate and scalable method for interpreting radiology reports and identifying PEs from ED data. If IT resources are available, this is a cost-effective approach to quality metric derivation for diagnostic processes in large health systems.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
Similar papers in this journal
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
Similar papers in this journal
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 94%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 93%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.