Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data
Yang, Y.; Gu, B.; Hathaway, D. B.; Wyss, R.; Marengo, L.; Gibbons, J. B.; Lyndon, S.; Wu, J.; Chen, Q.; Liu, N.; Wang, P. S.; Celi, L. A.; Bates, D. W.; Lin, J.; Zhou, L.; Yang, J.
Show abstract
Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has focused on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to reproduce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Open-source models and reasoning models showed statistically higher stigma rates than proprietary (1.97% vs. 1.60%; p < 0.01) and non-reasoning models (2.35% vs. 1.70%; p < 0.0001), while the stigma rate difference between the general and medical models is not statistically significant (2.00% vs. 1.80%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.304; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of model-task pairs amplifying stigma in the original input notes. Applying prompt engineering as a destigmatizing approach helped reduce model stigma rates by as much as 91.91% without affecting the model performance. This study shows that stigmatizing language generation is common but reducible during LLMs' reasoning traces, suggesting that well-implemented approaches for LLM monitoring and destigmatizing will be essential for healthcare systems to implement.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 92%
Similar papers in this journal
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 93%
- Detecting Goals of Care Conversations in Clinical Notes with Active Learning 92%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 92%
Similar papers in this journal
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 94%
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 91%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.