Mitigating Automation Bias in Physician-LLM Diagnostic Reasoning Using Behavioral Nudges: A Randomized Controlled Trial
Qazi, I. A.; Ali, A.; Khawaja, A. U.; Akhtar, M. J.; Sheikh, A. Z.; Alizai, M. H.
Show abstract
As large language models (LLMs) enter clinical workflows, automation bias, the uncritical acceptance of automated output, poses a patient-safety risk. Optimal physician-AI collaboration requires trust calibration, matching scrutiny to LLM recommendation accuracy. We report a randomized trial evaluating a behavioral nudge to mitigate automation bias. Seventy-two AI-trained physicians were randomized to evaluate six vignettes alongside ChatGPT-5.1 recommendations, consulted at each physician's discretion; three contained deliberate, clinically significant errors. The treatment arm received a dual-component nudge: an anchoring cue reporting ChatGPT's benchmark accuracy to calibrate expectations, and a case-specific, selective-attention cue; a numeric accuracy rating and color-coded traffic light, derived from the mean of three distinct-family LLMs. The control group saw recommendations alone; blinded reviewers scored diagnostic reasoning against an expert rubric. The treatment group scored significantly higher (mean difference, 7.6 percentage-points; 95% CI, 1.4-13.9; P=0.016) than the control, suggesting a scalable strategy to preserve clinical judgment in LLM-assisted care. ClinicalTrials.gov registration: NCT07328815.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 92%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 92%
Similar papers in this journal
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 91%
Similar papers in this journal
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 91%
- Outcomes of a Smartphone-based Application with Live Health-Coaching Post-Percutaneous Coronary Intervention 91%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 91%
Similar papers in this journal
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 92%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 91%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.