How Large Language Models Can Affect Clinical Reasoning: A Randomized Clinical Trial
Levels, M.; Rounding, N.; Arif, L. S.; Berg, J.; Cals, J. W.; de Boer, D.; de Bont, E. G.; Dijksman, S.; Findyartini, A.; Fouarge, D.; Gmyrek, P.; Greviana, N.; Leijenaar, R.; Manji, S.; Mbithi, A.; Obungu, N.; Pujitresnani, A.; Rianga, R.; Soemantri, D.; Sokwalla, S. M. R.; Steens, S.; Velasco, L.; Wildan, A.; Yusuf, P. A.; Freqin, M.-C.
Show abstract
ImportanceLLMs have encoded a vast array of medical knowledge and are being integrated into clinical settings as decision-support tools to improve physician performance across various aspects of care. However, evidence of the impact of LLMs on the clinical reasoning of physicians remains limited. ObjectiveTo evaluate the impact of LLM on core aspects of physicians clinical reasoning: diagnostic reasoning, information gathering, and management reasoning in primary care scenarios. Design, Setting, and ParticipantsWe conducted three identical randomized controlled trials (RCTs) in 2024-2025 with 249 physicians in Indonesia, Kenya, and the Netherlands. Participants completed four or five clinical vignettes designed to simulate real-world primary care consultations, with half randomized to have access to ChatGPT-4o. Main Outcomes and MeasuresPhysician quality of care was evaluated using a rubric based on evidence-based clinical guidelines, scored across nine steps of the clinical reasoning process. Primary outcomes were quality scores for diagnostic reasoning, information gathering, and management. Secondary outcomes were quality per answer, number of answers, and less obvious answers. ResultsAccess to LLMs enhanced information gathering and management reasoning across all countries. Physicians who were assigned the LLM achieved significantly better quality-of-care scores in diagnostic steps in Indonesia (b=7.9%, CI: 4.0% to 11.8%, p<.001) and Kenya (b= 15.1%, CI: 10.2% to 19.9%, p<.001) but not the Netherlands (b=1.4%, CI: -1.6% to 4.4%, p=1.00). Physicians with LLM access also performed better in investigative steps, in Indonesia (b=10.7%, CI: 4.2% to 17.1%, p=.004), Kenya (b=17.1%, CI: 10.3% to 23.9%, p<.001), and the Netherlands (b=11.9%, CI: 7.7% to 16.1%, p<.001). We also found LLM access affected physicians scores in management steps (Indonesia: b=15.7%, CI: 8.6% to 22.9%, p<.001; Kenya: b=27.3%, CI: 19.9% to 34.7% p<.001; the Netherlands: b=12.3%, CI: 7.1% to 17.5%, p<.001). We found that LLM access was less useful in management reasoning for more cognitively demanding cases compared to standard patient cases in Indonesia (b=-14.1%, CI: -21.4% to -6.8%, p<.001) and Kenya (b=-12.1%, CI: -19.6% to -4.6%, p=.006). Conclusions and RelevanceIn this cross-country randomized control trial, we assessed that access to an LLM had significant positive effects on physicians clinical reasoning. The effects we found are promising for the further roll-out of LLMs to supplement physicians in their care tasks. They also suggest that the extent to which LLMs can supplement physicians is context dependent. Key PointsO_ST_ABSQuestionC_ST_ABSTo what extent do large language models (LLMs) increase physicians quality of diagnostic reasoning, information gathering and management reasoning? FindingsIn a randomized clinical trial including 249 physicians in Indonesia, Kenya, and the Netherlands, access to an LLM significantly enhanced clinical reasoning performance in information gathering and management reasoning across all countries, and diagnostic reasoning in Kenya and Indonesia. MeaningThis study shows that the use of an LLM can enhance clinical reasoning of physicians. Further research is needed to effectively understand the augmentation of physician clinical practice.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 95%
- How and why do Quality Circles work for General Practitioners - a realist approach 95%
- Physicians in the management and leadership of health care: A systematic review of the conditions conducive to organizational performance 93%
Similar papers in this journal
- Development of the Tool for Advancing Practice Performance, a practice-level survey to assess primary care structures and processes 94%
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
Similar papers in this journal
- ‘Leading from the front’ implementation strategies increase the success of influenza vaccination drives among healthcare workers: A reanalysis of Systematic Review evidence using Intervention Component Analysis (ICA) and Qualitative Comparative Analysis (QCA) 92%
- The Monash Learning Health System Maturity Matrix: Codesign of a Tool to Measure and Guide Improvement in Complex Health System Behaviour 91%
- Dissatisfaction with family members’ medical care: the relationship with trust in personal physicians and physicians generally in Japan 91%
Similar papers in this journal
- Interventions To Improve Patient Safety During The COVID-19 Pandemic: A Systematic Review 92%
- Racial and Ethnic Diversity in Clinical Studies Reported to ClinicalTrials.gov, 2009-2024 90%
- How, Why And Under What Circumstances Does A Quality Improvement Collaborative Build Knowledge And Skills In Clinicians Working With People With Dementia? A Realist Informed Process Evaluation 90%
Similar papers in this journal
- Development of a consensus extension of the estimands framework for cluster randomised trials (CRT-estimands): results from an international Delphi study 91%
- Operational Strategies among Infectious Disease Clinical Trial Sites During Pandemics: A Scoping Review 91%
- ReachUHC: A randomized controlled trial study protocol of a mobile phone-based reminder and automatic renewal intervention to increase health insurance renewal rates in Kumasi, Ghana 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.