Evaluating the Influence of Demographic Identity in the Medical Use of Large Language Models
Lee, S.; Cho, W. I.; Lee, Y.; Park, C.; Park, C.; Ko, T.
Show abstract
As large language models (LLMs) are increasingly adopted in medical decision-making, concerns about demographic biases in AIgenerated recommendations remain unaddressed. In this study, we systematically investigate how demographic attributes--specifically race and gender--affect the diagnostic, medication, and treatment decisions of LLMs. Using the MedQA dataset, we construct a controlled evaluation framework comprising 20,000 test cases with systematically varied doctor-patient demographic pairings. We evaluate two LLMs of different scales: Claude 3.5 Sonnet, a highperformance proprietary model, and Llama 3.1-8B, a smaller open-source alternative. Our analysis reveals significant disparities in both accuracy and bias patterns across models and tasks. While Claude 3.5 Sonnet demonstrates higher overall accuracy and more stable predictions, Llama 3.1-8B exhibits greater sensitivity to demographic attributes, particularly in diagnostic reasoning. Notably, we observe the largest accuracy drop when Hispanic patients are treated by White male doctors, underscoring potential risks of bias amplification. These findings highlight the need for rigorous fairness assessments in medical AI and inform strategies to mitigate demographic biases in LLM-driven healthcare applications.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
Similar papers in this journal
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
Similar papers in this journal
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 93%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 92%
- User Experience Evaluation of Cogscreen for Screening Mild Cognitive Impairment: Formative and Summative Evaluation 91%
Similar papers in this journal
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 94%
- Clinical code sets and the problem of redundancy in code set repositories 92%
- Robust Disease Prognosis via Diagnostic Knowledge Preservation: A Sequential Learning Approach 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.