Influence of a Large Language Model on Diagnostic Reasoning: A Randomized Clinical Vignette Study
Goh, E.; Gallo, R.; Hom, J.; Strong, E.; Weng, Y.; Kerman, H.; Cool, J.; Kanjee-Khoja, Z.; Parsons, A. S.; Ahuja, N.; Horvitz, E.; Olson, A. P.; Rodman, A.; Yang, D.; Milstein, A.; Chen, J. H.
Show abstract
ImportanceDiagnostic errors are common and cause significant morbidity. Large language models (LLMs) have shown promise in their performance on both multiple-choice and open-ended medical reasoning examinations, but it remains unknown whether the use of such tools improves diagnostic reasoning. ObjectiveTo assess the impact of the GPT-4 LLM on physicians diagnostic reasoning compared to conventional resources. DesignMulti-center, randomized clinical vignette study. SettingThe study was conducted using remote video conferencing with physicians across the country and in-person participation across multiple academic medical institutions. ParticipantsResident and attending physicians with training in family medicine, internal medicine, or emergency medicine. Intervention(s)Participants were randomized to access GPT-4 in addition to conventional diagnostic resources or to just conventional resources. They were allocated 60 minutes to review up to six clinical vignettes adapted from established diagnostic reasoning exams. Main Outcome(s) and Measure(s)The primary outcome was diagnostic performance based on differential diagnosis accuracy, appropriateness of supporting and opposing factors, and next diagnostic evaluation steps. Secondary outcomes included time spent per case and final diagnosis. Results50 physicians (26 attendings, 24 residents) participated, with an average of 5.2 cases completed per participant. The median diagnostic reasoning score per case was 76.3 percent (IQR 65.8 to 86.8) for the GPT-4 group and 73.7 percent (IQR 63.2 to 84.2) for the conventional resources group, with an adjusted difference of 1.6 percentage points (95% CI -4.4 to 7.6; p=0.60). The median time spent on cases for the GPT-4 group was 519 seconds (IQR 371 to 668 seconds), compared to 565 seconds (IQR 456 to 788 seconds) for the conventional resources group, with a time difference of -82 seconds (95% CI -195 to 31; p=0.20). GPT-4 alone scored 15.5 percentage points (95% CI 1.5 to 29, p=0.03) higher than the conventional resources group. Conclusions and RelevanceIn a clinical vignette-based study, the availability of GPT-4 to physicians as a diagnostic aid did not significantly improve clinical reasoning compared to conventional resources, although it may improve components of clinical reasoning such as efficiency. GPT-4 alone demonstrated higher performance than both physician groups, suggesting opportunities for further improvement in physician-AI collaboration in clinical practice.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 94%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 92%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 95%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 93%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 92%
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 95%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.