Prompt Engineering Enables Open-Source LLMs to Match Proprietary Models in Diagnostic Accuracy for Annotation of Radiology Reports
Petersen, L. A.; Beck, M. S.; Xu, J. J.; Andersen, M. B.; Bruun, F. J.
Show abstract
AimThe aim of this study was to test whether open-source Large Language Models (LLMs) can match the diagnostic accuracy of proprietary models in annotating Danish trauma radiology reports across three clinical findings. Materials and MethodsThis retrospective study included 2,939 radiology reports of trauma radiographs collected from three Danish emergency departments. The data were split, with 600 cases for prompt engineering and 2,339 for model evaluation. Eight LLMs, GPT-4o and GPT-4o-mini (OpenAI), and six Llama3 variants (Meta) were prompted to annotate the reports for fractures, effusions, and luxations. The reference standard was human annotations. The diagnostic performance was assessed using accuracy, sensitivity, specificity, PPV, and NPV with 95% confidence intervals. ResultsPrompt engineering improved the Match-score for Llama3-8b from 77.8% (95% CI: 74.4% - 81.1%) to 94.3% (95% CI: 92.5% - 96.2%). GPT-4o achieved the highest overall diagnostic accuracy at 97.9% (95% CI: 97.3% - 98.5%), followed by Llama3.1-405b (97.1% (95% CI: 96.4% - 97.8%)), GPT-4o-mini (96.9% (95% CI: 96.2% - 97.6%)), Llama3-8b (96.9% (95% CI: 95.9% - 97.3%)), and Llama3.1-70b (96.0% (95% CI: 95.2% - 96.8%)). Across the three specific findings, all models performed best for fractures, whereas effusion and luxation were more prone to errors. Of the error types, Semantic Confusion was the most frequent, with 53.2% to 59.4% of misclassifications. ConclusionSmall, open-source LLMs can accurately annotate Danish trauma radiology reports when supported by effective prompt engineering, achieving accuracy levels that rival proprietary competitors. They offer a viable, privacy-conscious alternative for clinical use, even in a low-resource language setting.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 96%
- Fully automatic segmentation of craniomaxillofacial CT scans for computer-assisted orthognathic surgery planning using the nnU-Net framework 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 94%
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 94%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 93%
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 93%
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 92%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 92%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 92%
- Tracking And Predicting COVID-19 Radiological Trajectory Using Deep Learning On Chest X-Rays: Initial Accuracy Testing 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.