Impact of LLM Assistance on Physician Decision Making: A Multi-Country Randomized Controlled Trial
Rounding, N.; Arif, L. S.; Berg, J.; Cals, J. W.; de Boer, D.; de Bont, E. G.; Dijksman, S.; Findyartini, A.; Fouarge, D.; Freqin, M.-C.; Gmyrek, P.; Greviana, N.; Leijenaar, R.; Manji, S.; Mbithi, A.; Obungu, N.; Pujitresnani, A.; Rianga, R.; Soemantri, D.; Sokwalla, S. M. R.; Steens, S.; Velasco, L.; Wildan, A.; Yusuf, P. A.; Levels, M.
Show abstract
Disparities in the quality of healthcare persist globally, with poor-quality care contributing significantly to preventable mortality, particularly in low- and middle-income countries. While digital technologies, including generative artificial intelligence (AI), hold promise for improving clinical decision-making, their global effectiveness and potential to mitigate cross-country variation remain underexplored. We conducted a parallel-group randomized controlled trial across three economically diverse countries--Indonesia, Kenya, and the Netherlands--to evaluate the impact of large language model (LLM) access on physician performance using standardized clinical vignettes. Physicians (N=249) were randomly assigned to either a control group or an intervention group with access to GPT-4o. Results showed that LLM access significantly improved clinical performance, with the largest effect in Kenya (18%, 95% CI: 12.7 to 23.2, p<0.001), followed by Indonesia (10.7%, 95% CI: 5.7 to 15.7, p<0.001) and the Netherlands (7.2%, 95% CI: 3.7 to 10.7, p<0.001). Notably, LLM access reduced cross-country performance disparities, particularly between Kenya and the Netherlands. However, distributional effects varied, with increased score dispersion in Indonesia and reduced variation in Kenya. Higher LLM usage was associated with greater performance gains, though some physicians without access outperformed those with access, suggesting that effective use depends on individual engagement. Our findings demonstrate that LLMs can enhance clinical performance across diverse settings while potentially narrowing global inequalities in care quality. Further research should explore mechanisms of effective LLM integration and long-term impacts on real-world clinical practice.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Spatio-temporal modelling of referrals to outpatient respiratory clinics in the integrated care system of the Morecambe Bay area, England 92%
- ‘Leading from the front’ implementation strategies increase the success of influenza vaccination drives among healthcare workers: A reanalysis of Systematic Review evidence using Intervention Component Analysis (ICA) and Qualitative Comparative Analysis (QCA) 92%
- Dissatisfaction with family members’ medical care: the relationship with trust in personal physicians and physicians generally in Japan 92%
Similar papers in this journal
- Development of the Tool for Advancing Practice Performance, a practice-level survey to assess primary care structures and processes 93%
- Essential Indicators of Quality in Primary Care Settings: An Evidence-Based, Structured, Expert Approach 93%
- COVI-Prim survey: Challenges for Austrian and German general practitioners during initial phase of COVID-19 93%
Similar papers in this journal
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 94%
- How and why do Quality Circles work for General Practitioners - a realist approach 93%
- Physicians in the management and leadership of health care: A systematic review of the conditions conducive to organizational performance 93%
Similar papers in this journal
- Factors influencing focused practice: A qualitative study of resident and early-career family physician practice choices 92%
- OpenSAFELY NHS Service Restoration Observatory 1: describing trends and variation in primary care clinical activity for 23.3 million patients in England during the first wave of COVID-19 92%
- Changes in General Practice use and costs with COVID-19 and telehealth initiatives 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.