Physician epistemic framing alters the accuracy of large language models for medical second opinions
Reis, F.; Kunde, W.; Balzer, F.; Boie, S. D.
Show abstract
Large language models (LLMs) are increasingly being explored as tools for medical second opinions, yet their performance is often evaluated under neutral benchmark conditions that may not reflect how clinicians actually query these systems. We investigated whether physician epistemic framing alters LLM accuracy when the underlying clinical evidence remains identical. In this preregistered factorial prompting study, three state-of-the-art LLMs were evaluated on 499 MedQA-derived clinical cases across five within-case request conditions: neutral baseline, confirmation-seeking with correct or incorrect physician hypotheses, and contradiction-seeking in which the physician expressed doubt about correct or incorrect hypotheses. Across 7,485 model responses, baseline accuracy was 93.79%. Accuracy remained similar when physicians sought confirmation of correct or incorrect hypotheses (93.65% and 93.72%, respectively), but declined when physicians expressed doubt about the correct answer (88.51%; odds ratio versus baseline 0.51, 95% CI 0.43-0.61; Holm-adjusted p < 0.001). Exact adoption of an incorrect physician hypothesis occurred in 28 of 1,497 confirmation-seeking responses, whereas 86 responses changed from correct at baseline to incorrect when physicians expressed doubt about the correct hypothesis. These failures were concentrated in ambiguous cases and were sometimes accompanied by high self-reported confidence. Our findings show that medical LLM accuracy is interaction-sensitive: the same clinical evidence can yield different outputs depending solely on how a second-opinion request is framed. Evaluations of clinical LLMs should therefore move beyond neutral benchmark accuracy and incorporate interaction-based scenarios that reflect real clinician use.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 91%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 91%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Evaluating algorithmic fairness in the presence of clinical guidelines: the case of atherosclerotic cardiovascular disease risk estimation 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.