Bias Amplification in Intersectional Subpopulations for Clinical Phenotyping by Large Language Models
Pal, R.; Garg, H.; Patel, S.; Sethi, T.
Show abstract
Large Language Models (LLMs) have demonstrated remarkable performance across diverse clinical tasks. However, there is growing concern that LLMs may amplify human bias and reduce performance quality for vulnerable subpopulations. Therefore, it is critical to investigate algorithmic underdiagnosis in clinical notes, which represent a key source of information for disease diagnosis and treatment. This study examines prevalence of bias in two datasets - smoking and obesity - for clinical phenotyping. Our results demonstrate that state-of-the-art language models selectively and consistently underdiagnosed vulnerable intersectional subpopulations such as young-aged-males for smoking and middle-aged-females for obesity. Deployment of LLMs with such biases risks skewing clinicians decision-making which may lead to inequitable access to healthcare. These findings emphasize the need for careful evaluation of LLMs in clinical practice and highlight the potential ethical implications of deploying such systems in disease diagnosis and prognosis.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 95%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 95%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 93%
- Toward Using Twitter Data to Monitor Covid-19 Vaccine Safety in Pregnancy 91%
- A Web-based, Mobile Responsive Application to Screen Healthcare Workers for COVID Symptoms: Descriptive Study 90%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 96%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 95%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.