The Effect of LLM Assistance on Diagnostic Accuracy: A Meta-Analysis
Tretow, I.; Schwebel, M.; Feuerriegel, S.; Treffers, T.; Welpe, I. M.
Show abstract
Large language models (LLMs) are increasingly used in clinical settings, yet their effect on diagnostic accuracy of physicians has not been systematically quantified. We conducted a systematic review and meta-analysis of studies analyzing LLM-assisted diagnosis published between January 2020 and June 2025. Across 15 studies (43 effect sizes; 498 physicians; 7,274 case evaluations), LLM assistance significantly improved diagnostic accuracy compared to physicians without LLM support (Hedges g = 0.20, 95% CI 0.12-0.29; P < .001). Although improvements were observed across multiple LLMs (e.g., GPT-4, AMIE, MedFound-DX-PA), medical fields (general medicine, radiology), and career stages of physicians (residents and attendings), the magnitude of the benefit varied substantially. These findings show that LLMs can improve diagnostic accuracy of physicians, but conditions for successful LLM assistance remain unclear. Further clinical evidence is needed to guide safe and effective integration into practice.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 92%
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 92%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 92%
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 93%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 92%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 92%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 91%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.