Automation Bias in Large Language Model Assisted Diagnostic Reasoning Among AI-Trained Physicians
Qazi, I. A.; Ali, A.; Khawaja, A. U.; Akhtar, M. J.; Sheikh, A. Z.; Alizai, M. H.
Show abstract
ImportanceLarge language models (LLMs) show promise for improving clinical reasoning, but they also risk inducing automation bias, an over-reliance that can degrade diagnostic accuracy. Whether AI-trained physicians are vulnerable to this bias when LLM use is voluntary remains unknown. ObjectiveTo determine whether exposure to erroneous LLM recommendations degrades AI-trained physicians diagnostic performance compared to error-free AI advice. DesignA single-blind randomized clinical trial was conducted from June 20 to August 15, 2025. SettingPhysicians were recruited from multiple medical institutions in Pakistan, participating through in-person or remote video conferencing. ParticipantsPhysicians registered with the Pakistan Medical and Dental Council with MBBS degrees, who had completed a 20-hour AI-literacy training covering LLM capabilities, prompt engineering, and critical evaluation of AI output. InterventionParticipants were randomized 1:1 to diagnose 6 clinical vignettes in 75 minutes. The control group received unmodified ChatGPT-4os diagnostic recommendations; the treatment groups recommendations contained deliberate errors in 3 of 6 vignettes. Physicians could voluntarily consult offered ChatGPT-4o recommendations alongside conventional diagnostic resources based on their clinical judgment. Main Outcomes and MeasuresPrimary outcome was the diagnostic reasoning accuracy (percentage), assessed by three blinded physicians using an expert-validated rubric to evaluate: differential diagnosis accuracy, appropriateness of supporting and opposing evidence, and quality of recommended diagnostic steps. Secondary outcome was the top-choice diagnosis accuracy. ResultsForty-four physicians (22 treatment, 22 control) participated. Physicians receiving error-free recommendations achieved mean (SD) diagnostic accuracy of 84.9% (19.7%), whereas those exposed to flawed recommendations scored 73.3% (30.5%), resulting in an adjusted mean difference of -14.0 percentage points (95% CI: -8.3 to -19.7; P <.0001). Top-choice diagnosis accuracy per case was 76.1% (42.5) in the treatment group and 90.5% (28.9) in the control group, with an adjusted difference of -18.3 percentage points (95% CI, -26.6 to -10.0; P <.0001). Conclusions and RelevanceThis trial demonstrates that erroneous LLM recommendations significantly degrade physicians diagnostic performance by inducing automation bias, even in AI-trained physicians. Voluntary deference to flawed AI output highlights critical patient safety risk, necessitating robust safeguards to ensure human oversight before widespread clinical deployment. Trial RegistrationClinicalTrials.gov Identifier: NCT06963957
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 92%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Observer: Creation of a Novel Multimodal Dataset for Outpatient Care Research 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 95%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 91%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 91%
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 92%
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 91%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 91%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 93%
- Clinical practice guidelines and recommendations in the context of the COVID-19 pandemic: systematic review and critical appraisal 92%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.