Large Language Model Influence on Management Reasoning: A Randomized Controlled Trial
Ethan, E.; Gallo, R.; Strong, E.; Weng, Y.; Kerman, H.; Freed, J.; Cool, J. A.; Kanjee, Z.; Lane, K.; Parsons, A. S.; Ahuja, N.; Horvitz, E.; Yang, D.; Milstein, A.; Olson, A. P.; Hom, J.; Chen, J. H.; Rodman, A.
Show abstract
ImportanceLarge language model (LLM) artificial intelligence (AI) systems have shown promise in diagnostic reasoning, but their utility in management reasoning with no clear right answers is unknown. ObjectiveTo determine whether LLM assistance improves physician performance on open-ended management reasoning tasks compared to conventional resources. DesignProspective, randomized controlled trial conducted from 30 November 2023 to 21 April 2024. SettingMulti-institutional study from Stanford University, Beth Israel Deaconess Medical Center, and the University of Virginia involving physicians from across the United States. Participants92 practicing attending physicians and residents with training in internal medicine, family medicine, or emergency medicine. InterventionFive expert-developed clinical case vignettes were presented with multiple open-ended management questions and scoring rubrics created through a Delphi process. Physicians were randomized to use either GPT-4 via ChatGPT Plus in addition to conventional resources (e.g., UpToDate, Google), or conventional resources alone. Main Outcomes and MeasuresThe primary outcome was difference in total score between groups on expert-developed scoring rubrics. Secondary outcomes included domain-specific scores and time spent per case. ResultsPhysicians using the LLM scored higher compared to those using conventional resources (mean difference 6.5 %, 95% CI 2.7-10.2, p<0.001). Significant improvements were seen in management decisions (6.1%, 95% CI 2.5-9.7, p=0.001), diagnostic decisions (12.1%, 95% CI 3.1-21.0, p=0.009), and case-specific (6.2%, 95% CI 2.4-9.9, p=0.002) domains. GPT-4 users spent more time per case (mean difference 119.3 seconds, 95% CI 17.4-221.2, p=0.02). There was no significant difference between GPT-4-augmented physicians and GPT-4 alone (-0.9%, 95% CI -9.0 to 7.2, p=0.8). Conclusions and RelevanceLLM assistance improved physician management reasoning compared to conventional resources, with particular gains in contextual and patient-specific decision-making. These findings indicate that LLMs can augment management decision-making in complex cases. Trial RegistrationClinicalTrials.gov Identifier: NCT06208423; https://classic.clinicaltrials.gov/ct2/show/NCT06208423 Key PointsO_ST_ABSQuestionC_ST_ABSDoes large language model (LLM) assistance improve physician performance on complex management reasoning tasks compared to conventional resources? FindingsIn this randomized controlled trial of 92 physicians, participants using GPT-4 achieved higher scores on management reasoning compared to those using conventional resources (e.g., UpToDate). MeaningLLM assistance enhances physician management reasoning performance in complex cases with no clear right answers.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Observer: Creation of a Novel Multimodal Dataset for Outpatient Care Research 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 94%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 94%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 91%
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 93%
Similar papers in this journal
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
- Development of the Tool for Advancing Practice Performance, a practice-level survey to assess primary care structures and processes 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 92%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.