Socio-Demographic Biases in Medical Decision-Making by Large Language Models: A Large-Scale Multi-Model Analysis
Omar, M.; Soffer, S.; Agbareia, R.; Bragazzi, N. L.; Apakama, D. U.; Horowitz, C. R.; Charney, A.; Freeman, R.; Kummer, B.; Glicksberg, B. S.; Nadkarni, G.; Klang, E.
Show abstract
Large language models (LLMs) are increasingly integrated into healthcare but concerns about potential socio-demographic biases persist. We aimed to assess biases in decision-making by evaluating LLMs responses to clinical scenarios across varied socio-demographic profiles. We utilized 500 emergency department vignettes, each representing the same clinical scenario with differing socio-demographic identifiers across 23 groups--including gender identity, race/ethnicity, socioeconomic status, and sexual orientation--and a control version without socio-demographic identifiers. We then used Nine LLMs (8 open source and 1 proprietary) to answer clinical questions regarding triage priority, further testing, treatment approach, and mental health assessment, resulting in 432,000 total responses. We performed statistical analyses to evaluate biases across socio-demographic groups, with results normalized and compared to control groups. We find that marginalized groups--including Black, unhoused, and LGBTQIA+ individuals--are more likely to receive recommendations for urgent care, invasive procedures, or mental health assessments compared to the control group (p < 0.05 for all comparisons). High-income patients were more often recommended advanced diagnostic tests such as CT scans or MRI, while low-income patients were more frequently advised to undergo no further testing. We observed significant biases across all models, both proprietary and open source regardless of the models size. The most pronounced biases emerged in mental health assessment recommendations. LLMs used in medical decision-making exhibit significant biases in clinical recommendations, perpetuating existing healthcare disparities. Neither model type nor size affects these biases. These findings underscore the need for careful evaluation, monitoring, and mitigation of biases in LLMs to ensure equitable patient care.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gender-affirming care, mental health, and economic stability in the time of COVID-19: a global cross-sectional study of transgender and non-binary people 95%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 94%
- Psychosocial factors associated with mental health and quality of life during the COVID-19 pandemic among low-income urban dwellers in Peninsular Malaysia 93%
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 94%
- Patients with affective disorders profit most from telemedical treatment: Evidence from a naturalistic patient cohort during the COVID-19 pandemic 92%
- Longitudinal trends and risk factors for depressed mood among Canadian adults during the first wave of COVID-19 92%
Similar papers in this journal
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 92%
- Virtual and remote opioid poisoning education and naloxone distribution programs: a scoping review 91%
- Defining Destigmatizing Design Guidelines for Use in Sexual Health-Related Digital Technologies: A Delphi Study 91%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 94%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
- Remote working in mental health services: a rapid umbrella review of pre-COVID-19 literature 93%
Similar papers in this journal
- Patterns of SARS-CoV-2 testing preferences in a national cohort in the United States 93%
- Belief in Conspiracy Theory about COVID-19 Predicts Mental Health and Well-being -- A Study of Healthcare Staff in Ecuador 93%
- Testing, Testing: What SARS-CoV-2 testing services do adults in the United States actually want? 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.