The impact of large language models on diagnostic reasoning among LLM-trained physicians: a randomized clinical trial
Qazi, I. A.; Ali, A.; Khawaja, A. U.; Akhtar, M. J.; Sheikh, A. Z.; Alizai, M. H.
Show abstract
Diagnostic errors remain a pervasive yet preventable source of patient harm, with resourcelimited healthcare systems in low- and middle-income countries (LMICs) facing disproportionately higher diagnostic error rates due to limited access to diagnostic tools, specialists, and decision support systems. Large language models (LLMs) offer potential to bridge diagnostic gaps in these settings but can generate inaccurate information, making comprehensive AI-literacy training for physicians essential before deployment. However, whether structured AI-literacy training translates into improved diagnostic reasoning remains unknown. We conducted a single-blind randomized controlled trial involving 60 licensed physicians from multiple medical institutions in Pakistan, a LMIC, between January 10, 2025, and May 17, 2025. Participants completed a novel 20-hour AI-literacy curriculum covering LLM capabilities, limitations, and appropriate use. Post-training, physicians were randomized to either LLM access plus conventional resources or conventional resources only, with 75 minutes allocated to review up to 6 clinical vignettes. The primary outcome was diagnostic reasoning score (percentage) from a validated, expert-graded rubric assessing differential diagnosis, supporting/opposing factor appropriateness, and next steps; secondary outcome was time per vignette (seconds). Of 58 physicians completing the study, those with LLM access achieved mean diagnostic reasoning scores of 71.4% versus 42.6% with conventional medical resources alone, yielding an adjusted difference of 27.5 percentage points (95% CI, 22.8 to 32.2; P < 0.001). Mean time per case was similar between groups (603.8 vs. 635 seconds; adjusted difference -6.4 seconds, 95% CI -68.2 to 55.3; P = 0.84). While LLM alone outperformed the trained physician group by 11.5 percentage points (95% CI, 5.5 to 17.5; P < 0.001), in 38% of cases the physician plus LLM group surpassed the median LLM-alone performance, highlighting physician-AI complementarities. This trial suggests that AI-literacy training can enable physicians in resourcelimited settings to effectively leverage LLMs for enhanced diagnostic reasoning to address diagnostic gaps in LMICs (ClinicalTrials.gov: NCT06774612).
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 95%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 91%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 91%
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 92%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.