Certified large language model-based diagnostic decision support in rheumatology: the ALLIANCE multicentre randomised controlled trial
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 90%
Similar papers in this journal
- Risk of death among people with rare autoimmune diseases compared to the general population in England during the 2020 COVID-19 pandemic 90%
- Use of Physician Global Assessment (PGA) in Systemic lupus erythematosus: a systematic review of its psychometric properties 90%
- Characteristics, outcomes, and mortality amongst 133,589 patients with prevalent autoimmune diseases diagnosed with, and 48,418 hospitalised for COVID-19: a multinational distributed network cohort analysis 89%
Similar papers in this journal
- Incidence and management of inflammatory arthritis in England before and during the COVID-19 pandemic: a population-level cohort study using OpenSAFELY 92%
- Versus Arthritis Musculoskeletal Disorders Research Advisory Group Priority Setting Exercise Protocol 91%
- Risk of severe COVID-19 outcomes associated with immune-mediated inflammatory diseases and immune modifying therapies: a nationwide cohort study in the OpenSAFELY platform 88%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 91%
- Crowdfunding Medical Care: A Comparison of Online Medical Fundraising in Canada, the United Kingdom, and the United States 90%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 89%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 91%
- Feasibility trial of a new digital training package to enhance primary care practitioners' communication of clinical empathy and realistic optimism 91%
- Clinical academic research in the time of Corona: a simulation study in England and a call for action 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.