Large Language Models in Radiology Reporting - A Systematic Review of Performance, Limitations, and Clinical Implications
Artsi, Y.; Klang, E.; Collins, J. D.; Glicksberg, B. S.; Korfiatis, P.; Nadkarni, G.; Sorin, V.
Show abstract
BackgroundLarge language models (LLMs) have emerged as potential tools for automated radiology reporting. However, concerns regarding their fidelity, reliability, and clinical applicability remain. This systematic review examines the current literature on LLM-generated radiology reports. MethodsWe conducted a systematic search of MEDLINE, Google Scholar, Scopus, and Web of Science to identify studies published between January 2015 and February 2025. Studies evaluating LLM-generated radiology reports were included. The study follows PRISMA guidelines. Risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2) tool. ResultsNine studies met the inclusion criteria. Of these, six evaluated full radiology reports, while three focused on impression generation. Six studies assessed base LLMs, and three evaluated fine-tuned models. Fine-tuned models demonstrated better alignment with expert evaluations and achieved higher performance on natural language processing metrics compared to base models. All LLMs showed hallucinations, misdiagnoses, and inconsistencies. ConclusionLLMs show promise in radiology reporting. However, limitations in diagnostic accuracy and hallucinations necessitate human oversight. Future research should focus on improving evaluation frameworks, incorporating diverse datasets, and prospectively validating AI-generated reports in clinical workflows.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 93%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 96%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 93%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 91%
Similar papers in this journal
- Natural language inference for clinical registry curation 92%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 92%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 92%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- On evaluation metrics for medical applications of artificial intelligence 93%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 92%
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 94%
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.