Recognizing "Conformity Bias" in Large Language Models: A New Risk for Clinical Use
Nowroozzadeh, M. H.; Salari, R.
Show abstract
ObjectivesThe aim of the present study is to systematically investigate the phenomenon of Conformity Bias in contemporary LLMs, specifically evaluating how repeated probing with incorrect information influences model outputs in a clinical context. Methods4 LLMs including GPT-4o, Gemini-1.5 Flash, Claude-3 Haiku, and GPT-o1 were systematically evaluated through 20 clinical questions focused on ocular disease treatments. Standard queries were followed by probing questions suggesting incorrect treatments. Model responses were analyzed to assess the emergence of Conformity Bias and compared using chi-squared testing. ResultsCorrect response rates after successive probing questions were alarmingly low: 25% (GPT-4o), 10% (Gemini-1.5 Flash), 0% (Claude-3 Haiku), and 25% (GPT-o1) (P < 0.001). Across models, the tendency to conform to incorrect user suggestions increased with repeated probing. ConclusionConformity Bias represents a dynamic, user-induced vulnerability in LLMs, distinguishable from training-dependent biases. Its presence underscores the necessity for model designs resistant to misleading user interactions and emphasizes the importance of cross-verification with clinical guidelines. As healthcare systems increasingly integrate AI tools, understanding and mitigating Conformity Bias is imperative to protect patient safety and maintain clinical integrity.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 95%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 95%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 94%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 92%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 94%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.